The Real Cost of Running an AI Voice Agent in 2026

Understanding what drives AI voice pricing is essential for anyone building or scaling a production-grade agent. The AI voice agent cost in 2026 is shaped by several core components:

Core price components:

Cost breakdown for a production-grade AI voice agent

1. Speech Synthesis (Text-to-Speech / TTS)

What is TTS and why is it needed?

Text-to-Speech allows a voice agent to speak responses in a natural-sounding voice, which is essential for phone-based or voice interactions. Understanding the real-time vs turn-based TTS architecture is crucial for optimizing both costs and performance.

What are the key TTS considerations?

What are the usual TTS billing models?

What are the usual TTS price ranges?

Roughly $4 to $200 per 1 million characters in 2026. Standard cloud voices (Polly Standard, Google Standard) sit at ~$4/M. Mainstream neural voices (Polly Neural, Google Neural2, Azure Neural, OpenAI gpt-4o-mini-tts) cluster around $12–$22/M. Premium real-time models (Cartesia Sonic 3 ~$35/M effective, ElevenLabs Turbo/Flash v2.5 ~$50/M) and ultra-realistic flagships (ElevenLabs v3 ~$100/M, Hume Octave ~$50–$150/M, Google Studio $160/M) dominate the high end. Open-source self-hosted (Kokoro, CosyVoice, Qwen3-TTS) has no per-character rate at all: the bill is the box it runs on. Translates to roughly $0.005–$0.04 per minute of generated speech.

What are the key TTS providers?

2. Speech Recognition (Speech-to-Text / STT / ASR)

What is STT and why is it needed?

Speech-to-Text (STT), also known as Automatic Speech Recognition (ASR), is the component that understands what the user is saying. It transcribes the caller's speech into text so that it can be processed by an LLM. Without accurate STT, the voice agent cannot know the user's request. For an in-depth comparison of STT and TTS providers, check our comprehensive STT/TTS selection guide.

Key considerations?

Usual billing models:

Usual price ranges:

Streaming STT spans roughly $0.0015–$0.024 per minute in 2026. The cheapest end is self-hosted or specialized models (NVIDIA Parakeet via Together at $0.0015/min, Cartesia Ink-Whisper at $0.00217/min). Mainstream cloud providers cluster around $0.0025–$0.012/min. Realtime/multimodal models like OpenAI gpt-realtime bundle STT+LLM at ~$0.06/min. Volume commits can drop streaming rates by 30–50%.

Key providers

3. Large Language Models (LLMs)

What is it and why is it needed?

The LLM is essentially the "brain" of the voice agent. It takes the transcribed text input and processes language – understanding intent and context – then generates a text response. In a modern AI voice agent, an LLM (like GPT-style models) enables natural conversation, dynamic responses, and handling of free-form user input that rule-based systems cannot. For detailed guidance on selecting the right model, see our comprehensive LLM selection guide.

Key considerations:

Usual billing models:

Token-based billing – this is standard for most LLM APIs. You pay for input tokens (the text you send in, including conversation history) and output tokens (the text the model generates). A token is roughly 0.75 words, so 1,000 tokens ~ 750 words.

Key providers:

4. Voice Agent Platform

What is it and why is it needed?

A Voice Agent Platform is an orchestration layer or framework for building the actual voice bot/agent. It typically ties together the telephony, STT, LLM, and TTS components, handling the call flow logic, state management, and integration with any backend systems.

Key considerations:

Usual billing models:

Usage-based (per minute of call) – The platform might charge you per minute of voice call handled by the AI agent. This often encompasses the underlying costs (STT, TTS, etc.), essentially as a bundled rate.

Usual price ranges:

$0.01–$0.14 per minute in 2026. Pure orchestration platforms (Pipecat Cloud, LiveKit Cloud Agents) charge $0.01/min and pass model costs through at vendor cost. Mid-market bring-your-own-key platforms (Vapi $0.05, Retell $0.055, Synthflow $0.09) layer an explicit per-minute fee on raw component costs. Bundled platforms (Bland AI $0.11–$0.14) embed LLM/STT/TTS/telephony in a single rate, with an estimated 15–40% markup on the underlying components. Speech-to-speech bundles (Ultravox $0.05) include the model itself.

Key providers:

5. Transport (Telephony / WebRTC)

What is it and why is it needed?

Transport carries audio between user and agent. Two modes: PSTN telephony (phone calls over SIP trunks) and WebRTC (browser- or app-embedded calls). PSTN providers handle phone numbers, inbound/outbound call routing, and SIP session control. WebRTC providers handle media servers, NAT traversal, and adaptive bitrate. Pick telephony if users dial a number; pick WebRTC if they hit a button on your web or mobile app.

Usual billing models:

Usual price ranges:

Telephony US: $0.005–$0.02/min for inbound + outbound average, plus $0.50–$2/mo per number, plus regulatory surcharge. WebRTC audio-only: $0.0004–$0.004/min on managed tiers, $0 on self-host (excluding infra).

Key telephony providers:

Key WebRTC providers:

Summary: AI voice agent cost per minute

Component Typical cost per minute Notes
TTS $0.005–$0.04 Billed per character. OSS self-host = $0; ElevenLabs v3 / Studio voices top end.
STT $0.0015–$0.024 Billed per audio minute. NVIDIA Parakeet cheapest, Google Chirp / Azure top end.
LLM $0.005–$0.05 Per-token billing × ~1.8× reality factor for context growth, interrupts, tool calls.
Platform $0.01–$0.14 Pipecat / LiveKit $0.01; Vapi $0.05; Bland bundled $0.14 (incl. components).
Transport $0.0004–$0.02 WebRTC low end; PSTN telephony high end. Add 5–15% USF and DID rental.
Total $0.13–$0.30 Production-grade voice agents in 2026. Cost-effective stacks land $0.05–$0.10; premium / managed / Realtime-API stacks $0.30+.

Note: LLMs are priced per token, TTS per character, STT per minute, transport per minute or participant-minute. Per-minute estimates above assume typical voice usage (150 WPM, ~4 turns/min) and apply a 1.8× reality factor on LLM cost to capture conversation-history token growth (compounds O(n²) with turns), function-calling round-trips, and barge-in handling. Bundled platforms like Bland AI embed an estimated 15–40% markup on the underlying STT/TTS/LLM. Telephony figures exclude US 5–15% USF / regulatory pass-through and per-DID monthly rental.

All Model Pricing

A self-hosted row has no vendor rate. It requires hosting.

Large Language Models (LLM)

ModelProviderInput (1M tokens)Output (1M tokens)
Qwen3.6-FlashAlibaba$0.1875$1.13Pricing
Qwen3.6-PlusAlibaba$0.325$1.95Pricing
Qwen3.8 27BAlibaba$0.400$3.00Pricing
Qwen3.8-FlashAlibaba$0.150$0.470Pricing
Qwen3.8-MaxAlibaba$2.00$6.00Pricing
Claude Fable 5.1Anthropic$10.00$50.00Pricing
Claude Haiku 4.5Anthropic$1.10$5.50Pricing
Claude Opus 5Anthropic$5.50$27.50Pricing
Claude Opus 5.5Anthropic$4.00$20.00Pricing
Claude Sonnet 5Anthropic$2.00$10.00Pricing
DeepSeek V4.1 FlashDeepSeek$0.300$1.20Pricing
Gemini 3.1 Flash LiteGoogle$0.250$1.50Pricing
Gemini 3.5 Flash LiteGoogle$0.330$2.75Pricing
Gemini 3.8 FlashGoogle$1.50$7.50Pricing
Gemma 4 26B-A4BGoogle$0.100$0.340Pricing
Gemma 4 31BGoogle$0.750$1.00Pricing
Gemma 12B Q4 (self-hosted)Googlerequires hosting
Llama 4 MaverickMeta$0.350$1.15Pricing
Muse Spark 1.3Meta$1.25$4.25Pricing
Ministral 3 8BMistral$0.150$0.150Pricing
Mistral Large 3Mistral$0.500$1.50Pricing
Mistral Medium 3.5Mistral$1.50$7.50Pricing
Mistral Small 4Mistral$0.165$0.660Pricing
Nemotron 3.5 LightningNVIDIA$0.070$0.200Pricing
GPT-5.4OpenAI$2.50$15.00Pricing
GPT-5.4 miniOpenAI$0.750$4.50Pricing
GPT-5.4 nanoOpenAI$0.200$1.25Pricing
GPT-5.5OpenAI$5.50$33.00Pricing
GPT-5.6 LunaOpenAI$0.220$1.32Pricing
GPT-5.6 SolOpenAI$4.00$20.00Pricing
GPT-5.6 TerraOpenAI$2.20$13.20Pricing
GPT-6 AstraOpenAI$10.00$50.00Pricing
GPT-6 LunaOpenAI$0.110$0.550Pricing
GPT-6 SolOpenAI$2.20$11.00Pricing
gpt-oss-20bOpenAI$0.030$0.130Pricing
gpt-oss-120bOpenAI$0.050$0.250Pricing
PhoneLLM Alpha 1 (self-hosted)Pipecatrequires hostingPricing
Grok 4.3xAI$1.25$2.50Pricing
Grok 4.5xAI$2.00$6.00Pricing
Grok 4.6xAI$2.00$6.00Pricing
Grok 4.7xAI$1.60$4.80Pricing
GLM-5.2Z.ai$1.40$4.40Pricing
GLM-5.3Z.ai$1.40$4.40Pricing
GLM-5.3-FlashZ.ai$0.150$0.500Pricing
GLM-5.3-FlashXZ.ai$0.370$1.25Pricing

Text-to-Speech (TTS)

ModelProviderPrice (1K characters)
CosyVoice (self-hosted)Alibabarequires hosting
Qwen3-TTS 1.7B (self-hosted)Alibabarequires hostingPricing
Amazon Polly StandardAmazon$0.004Pricing
Cartesia Sonic 3Cartesia$0.050Pricing
Cartesia Sonic 3.5Cartesia$0.050Pricing
Cartesia Sonic 3.6Cartesia$0.050Pricing
XTTS v2 (self-hosted)Coquirequires hosting
Deepgram Aura-2Deepgram$0.030Pricing
ElevenLabs Flash v2.5ElevenLabs$0.050Pricing
ElevenLabs v3 ConversationalElevenLabs$0.050Pricing
Fish Audio S2.1 ProFish Audio$0.015Pricing
Google TTS Chirp 3 HDGoogle$0.030Pricing
Kokoro 82M (self-hosted)hexgradrequires hosting
Hume OctaveHume$0.150Pricing
Inworld TTS-2Inworld$0.025Pricing
Inworld TTS-2 FlashInworld$0.015Pricing
Azure AI Speech NeuralMicrosoft$0.015Pricing
MiniMax Speech 2.8 TurboMiniMax$0.060Pricing
Murf Falcon 2Murf$0.0131Pricing
Magpie-Multilingual 357M (self-hosted)NVIDIArequires hosting
OpenAI GPT-4o Mini TTSOpenAI$0.012Pricing
Chatterbox (self-hosted)Resemble AIrequires hosting
Rime CodaRime$0.050Pricing
Rime Mist v3Rime$0.030Pricing
Smallest AI Lightning V3.1Smallest AI$0.0175Pricing
Smallest AI Lightning V3.1 ProSmallest AI$0.0195Pricing
Soniox TTS RT v2Soniox$0.0153Pricing
Speechify Simba 3.0Speechify$0.010Pricing
Speechify Simba 3.2Speechify$0.010Pricing
StyleTTS 2 (self-hosted)StyleTTSrequires hosting
xAI Grok TTSxAI$0.015Pricing

Speech-to-Text (STT)

ModelProviderPrice (per minute)
AssemblyAI Universal-3.5 Pro RealtimeAssemblyAI$0.0075Pricing
AssemblyAI Universal-StreamingAssemblyAI$0.0025Pricing
Cartesia Ink-2Cartesia$0.009Pricing
Deepgram Flux (English)Deepgram$0.0077Pricing
Deepgram Flux MultilingualDeepgram$0.0078Pricing
Deepgram Nova-3Deepgram$0.0077Pricing
Deepgram Nova-3 MultilingualDeepgram$0.0092Pricing
ElevenLabs Scribe v2 RealtimeElevenLabs$0.0065Pricing
Gladia Solaria 1Gladia$0.0125Pricing
Google Speech-to-Text Chirp 2Google$0.016Pricing
Google Speech-to-Text Chirp 3Google$0.016Pricing
Inworld STT-1Inworld$0.0025Pricing
Azure Speech-to-Text StreamingMicrosoft$0.0167Pricing
Mistral Voxtral Mini Transcribe RealtimeMistral$0.006Pricing
Voxtral Mini (self-hosted)Mistralrequires hosting
Canary Qwen 2.5B (self-hosted)NVIDIArequires hosting
Nemotron ASR Streaming (self-hosted)NVIDIArequires hosting
Parakeet TDT 0.6B V3 (self-hosted)NVIDIArequires hosting
OpenAI gpt-4o-mini-transcribeOpenAI$0.003Pricing
OpenAI gpt-4o-transcribeOpenAI$0.006Pricing
OpenAI gpt-realtime-whisperOpenAI$0.017Pricing
Whisper Large v3 (self-hosted)OpenAIrequires hosting
Smallest AI PulseSmallest AI$0.004Pricing
Soniox STT RT v5Soniox$0.002Pricing
Speechmatics Linden 1Speechmatics$0.0027Pricing
NVIDIA Parakeet TDT 0.6B v3Together AI$0.0035Pricing
xAI Grok STTxAI$0.0033Pricing

Developer Platforms

PlatformProviderPrice (per minute)
No Platform (BYOK direct)Nonerequires hosting
Pipecat Cloud (agent-1x active)Daily$0.010Pricing
LiveKit (self-host)LiveKitrequires hostingPricing
LiveKit Cloud AgentsLiveKit$0.010Pricing
Millis AIMillis AI$0.020Pricing
Pipecat (self-host)Pipecatrequires hostingPricing
Retell AIRetell AI$0.055Pricing
SynthflowSynthflow$0.090Pricing
VapiVapi$0.050Pricing

Transport (Telephony / WebRTC)

ProviderModePrice (per minute)Transfer Fee
No Transportweb$0.000–
Bandwidthtelephony$0.0078–Pricing
Plivotelephony$0.0055–Pricing
SignalWiretelephony$0.0073–Pricing
Telnyxtelephony$0.004$0.100Pricing
Twilio (inbound)telephony$0.0085–Pricing
Twilio (outbound)telephony$0.014–Pricing
Agora RTCweb$0.001–Pricing
Daily.co WebRTCweb$0.001–Pricing
LiveKit Cloud (Scale)web$0.0004–Pricing
Self-host (LiveKit / mediasoup)webrequires hosting–Pricing