Voice AI &
Speech Synthesis APIs
Model all-in per-minute costs for voice agents: Deepgram Nova-3 STT + ultra-low latency LLMs + Cartesia/ElevenLabs TTS + SIP trunking.
Provider Performance & 2026 Commercial Rate Directory
Deterministic unit pricing, context constraints, prompt cache read multipliers, and verified production SLAs.
| Provider & Engine | Tier / Architecture | Context / Payload | Verified 2026 Base Rate | Cache / Volume Rate | p50 Turnaround | Enterprise SLA |
|---|---|---|---|---|---|---|
|
Cartesia Sonic 2026
Cartesia AI
|
Sub-100ms Ultra-Fast TTS | PCM / Opus Stream | $0.0070 / 1,000 chars | $0.0050 / 1k (Sonic Turbo) | 95 ms TTFB | 99.9% |
|
ElevenLabs Turbo v2.5
ElevenLabs
|
High-Fidelity Expressive TTS | Ultra-Realistic | $0.0300 / 1,000 chars | $0.0200 / 1k (Flash 2.5) | 220 ms TTFB | 99.95% |
|
Deepgram Nova-3
Deepgram
|
Streaming Speech-to-Text | Real-Time WebSocket | $0.0043 / audio minute | $0.0036 / min (Commitment) | 220 ms Turnaround | 99.99% |
|
Deepgram Nova-3 Medical
Deepgram
|
Clinical Domain STT | HIPAA BAA | $0.0075 / audio minute | Custom Enterprise | 240 ms Turnaround | 99.99% |
|
OpenAI Whisper Large-v3
OpenAI
|
Batch Transcription Leader | 25MB audio file | $0.0060 / audio minute | Self-Host Open Weights | 550 ms Turnaround | 99.95% |
|
OpenAI TTS HD
OpenAI
|
High-Definition TTS | 4,096 chars | $0.0300 / 1,000 chars | $0.0150 / 1k (TTS Standard) | 350 ms TTFB | 99.95% |
|
AssemblyAI Universal-2
AssemblyAI
|
Speech Understanding Core | Streaming WebSocket | $0.0065 / audio minute | $0.0020 / min (Nano Tier) | 400 ms Turnaround | 99.9% |
|
PlayHT 2.0 Turbo
PlayHT
|
Conversational Low-Latency | Streaming Audio | $0.0250 / 1,000 chars | Enterprise Tiered | 260 ms TTFB | 99.9% |
|
Resemble AI Rapid
Resemble AI
|
Real-Time Voice Cloning | Custom Cloned Voice | $0.0100 / 1,000 chars | Volume discounts | 180 ms TTFB | 99.9% |
|
Speechmatics Flow
Speechmatics
|
Autonomous Speech Assistant | End-to-End Voice | $0.0070 / audio minute | Specialized Dialects | 320 ms Turnaround | 99.95% |
|
LMNT Aurora Ultra-Low
LMNT
|
Edge Voice Generation | Streaming Raw PCM | $0.0080 / 1,000 chars | Flat subscriptions | 75 ms TTFB | 99.9% |
|
Hume AI EVI
Hume AI
|
Empathic Voice Interface | Prosody & Tone Stream | $0.0120 / call minute | SDK Integrated | 380 ms Turnaround | 99.9% |
Technical Architecture & Bill Shock Prevention
The Voice Activity Detection (VAD) Billing Bleed
Many telephony bots keep an open audio stream running continuously during a 5-minute phone call, paying per-second STT rates for awkward background room noise, coughs, and dial tones. Implementing a client-side Silero VAD (Voice Activity Detection) filter ensures you only transmit audio frames when human speech is detected, dropping billable STT audio minutes by 45% to 60%.
Full-Duplex Interruption Packet Flooding
When a human user interrupts an AI voice agent mid-sentence, naive architectures continue generating and streaming the remaining TTS audio frames down the WebSocket, wasting money on synthesized characters the user will never hear. A production system must issue an immediate cancellation signal (abort controller) to the TTS provider within 50ms of user speech onset.
import asyncio
import os
from cartesia import AsyncCartesia
# Production Low-Latency Interruption Handler & Voice Streamer
async def stream_voice_response(text_stream, audio_ws):
cartesia = AsyncCartesia(api_key=os.getenv("CARTESIA_API_KEY"))
# Initialize low-latency sonic context
voice_ctx = cartesia.tts.websocket()
try:
async for chunk in text_stream:
# Send incremental text chunks as the LLM generates them
await voice_ctx.send(
model_id="sonic-english",
transcript=chunk,
voice_id="694f12bc-c4db-490f-b1e4-a42c0d52f952",
continue_=True
)
# Forward binary audio frames directly to phone/browser
async for audio_chunk in voice_ctx.receive():
await audio_ws.send_bytes(audio_chunk["audio"])
except asyncio.CancelledError:
# User interrupted: cancel remote generation immediately
print("User barge-in detected! Halting TTS synthesis stream...")
await voice_ctx.cancel()
raise
Interactive Regional Cost & Latency Simulator
Model your monthly operational expenditure and projected latency across deployment zones.
Frequently Asked Questions
Commonly evaluated trade-offs, contractual pitfalls, and latency optimization rules.