Lowest-Latency Inference APIs for Voice and Realtime Agents: A Time to First Token TTFT-First Benchmark
Time to first token (TTFT) is the metric groups use to decide an inference API for voice. It can also be the metric that misleads them. TTFT marks when technology begins; a text-to-speech mannequin can not converse till a full clause arrives. Between these two factors sits the distinction between an agent that feels conversational and one which will get interrupted. This piece benchmarks each layer of the voice stack together with LLM, speech-to-text, text-to-speech, and speech-to-speech.
Why TTFT Is the Right Entry Point and the Wrong Finish Line
A voice agent is a latency finances with a language mannequin inside it. Every stage spends milliseconds the consumer can hear.
Time to first token (TTFT) is the interval between sending an inference request and receiving the primary token again. IBM’s definition frames it because the second a system transitions from idle to visibly energetic.
For chat, TTFT is shut to the entire story. For voice, it’s one time period in a sum.
The motive is mechanical. A text-to-speech mannequin can not synthesize half a phrase. It wants an entire clause or sentence earlier than it produces audio. LiveEquipment calls the ensuing metric time-to-first-sentence (TTFS), and argues in its Gemma 4 deployment post that TTFS is what customers really really feel.
That offers you two knobs reasonably than one. TTFT controls when technology begins. Tokens per second controls how briskly the primary sentence completes. A supplier that wins one and loses the opposite won’t really feel quick.
The Latency Budget: What One Voice Turn Actually Costs
LiveEquipment’s voice agents overview breaks a flip into STT at roughly 100–200ms, LLM at 300–500ms with streaming, TTS at 100–200ms, and community at 50–150ms over WebRTC. It places the sensible end-to-end goal at 700ms to 1.2s.
Kwindla Hultman Kramer, co-creator of Pipecat, has suggested focusing on 800ms median voice-to-voice latency, with a looser 1,500ms acceptable for a proof of idea. His tough arithmetic splits that 4 methods at roughly 200ms every: transport and media processing, STT plus phrase endpointing, LLM inference, and TTS.
Daily’s earlier work on the fastest voice bot provides the human baseline. Typical human response time in dialog is round 500ms. Pauses past 800ms begin to really feel unnatural.
Daily’s February 2026 voice-agent LLM benchmark interprets that into an LLM requirement straight. Natural dialog wants voice-to-voice underneath 1,500ms, which works out to roughly 700ms of TTFT finances for a text-mode LLM inside a transcription-to-LLM-to-voice harness.
That 700ms quantity is the bar to maintain each supplier in opposition to.
How to Read a TTFT Benchmark Without Being Misled
Before the tables, 5 methodology information that change what the numbers imply:
1. Workload form dominates: Artificial Analysis modified its default workload in March 2026. The web site now experiences 10k enter token prompts reasonably than 1k. Longer prompts increase each TTFT and output pace. LiveEquipment argues that is closer to reality for voice, as a result of manufacturing brokers front-load coverage, persona, escalation guidelines, retrieved knowledge, and instrument schemas.
2. Server location is baked in: Artificial Analysis assessments from a digital machine in Google Cloud’s us-central1-a zone. It states plainly that TTFT contains community latency and could benefit or drawback suppliers primarily based on the place they serve.
3. Reasoning tokens depend: In the Artificial Analysis definition, TTFT for a reasoning mannequin is the primary reasoning token, not the primary reply token. Those are separate columns.
4. Measure from the receiving facet: Daily notes that model providers sometimes quote TTFT internal to their inference stacks. Daily measures from request ship to first usable token off the API.
5. Runs should not repeatable: Daily is blunt about this: TTFT varies considerably between benchmark runs, and suppliers change inference stacks and generally weights with out altering mannequin names.
Layer 1: LLM Time to First Token
Figures under are from the Artificial Analysis API providers leaderboard, retrieved August 30, 2026. The “first chunk” column is TTFT. Workload is 10k enter tokens, single immediate, median over 72 hours.
Lowest measured first-chunk latency
| Provider | Model | TTFT | Output pace |
|---|---|---|---|
| Baseten | gpt-oss-120b (excessive) | 0.23s | 266 tok/s |
| Baseten | gpt-oss-120b (low) | 0.24s | 271 tok/s |
| DeepInfra | Nemotron 3 Ultra | 0.28s | 371 tok/s |
| Cohere | North Mini Code | 0.32s | 104 tok/s |
| Cohere | Command A+ | 0.40s | 239 tok/s |
| Baseten | Inkling Small | 0.42s | 337 tok/s |
| Modular | Gemma 4 31B (NVFP4) | 0.44s | 243 tok/s |
| Nebius | GLM-5.3-Flash | 0.46s | 206 tok/s |
| Fireworks | Nemotron 3.5 Lightning | 0.46s | 501 tok/s |
| Together AI | Kimi K2.7 Code | 0.47s | 245 tok/s |
| Cerebras | gpt-oss-120b (excessive) | 0.49s | 1,697 tok/s |
The throughput lure
Silicon distributors optimize for a unique metric than voice brokers want.
| Provider | Model | TTFT | Output pace |
|---|---|---|---|
| Cerebras | gpt-oss-120b (excessive) | 0.49s | 1,697 tok/s |
| Celeris | Celeris-1 | 0.62s | 1,612 tok/s |
| Cerebras | Gemma 4 31B | 0.53s | 1,351 tok/s |
| Groq | gpt-oss-20b (excessive) | 0.82s | 957 tok/s |
| SambaNova | gpt-oss-120b (excessive) | 0.92s | 706 tok/s |
| Groq | gpt-oss-120b (low) | 0.69s | 473 tok/s |
| Inception | Mercury 2 | 3.07s | 770 tok/s |
Mercury 2 is the clearest illustration. It is a diffusion-based language mannequin, and it generates 770 tokens per second. Its first chunk arrives at 3.07s. That is 4 occasions all the LLM finances for a pure dialog.
Cerebras and Groq are a unique case. Their TTFT is respectable and their throughput is outstanding. For TTFS particularly, that mixture is robust, as a result of the sentence completes nearly instantly after the primary token lands.
Frontier and proprietary endpoints
| Provider | Model | TTFT | Output pace |
|---|---|---|---|
| Amazon Bedrock | GPT-5.6 Luna (non-reasoning) | 0.59s | 181 tok/s |
| Amazon Bedrock | GPT-5.6 Terra (non-reasoning) | 0.72s | 103 tok/s |
| OpenAI | GPT-5.6 Luna (non-reasoning) | 0.74s | 113 tok/s |
| Gemini 3.7 Flash (low), AI Studio | 0.84s | 315 tok/s | |
| Anthropic | Claude 4.5 Haiku (non-reasoning) | 0.84s | 82 tok/s |
| Amazon Bedrock | Nova Micro | 0.86s | 264 tok/s |
| Gemini 3.5 Flash (minimal), AI Studio | 0.90s | 202 tok/s | |
| OpenAI | GPT-5.6 Sol (non-reasoning) | 1.06s | 71 tok/s |
Note the identical mannequin on totally different hosts. GPT-5.6 Luna non-reasoning measures 0.59s on Amazon Bedrock and 0.74s on OpenAI’s personal API. Hosting and routing matter as a lot because the weights.
The vendor-measured outlier
LiveEquipment publishes TTFT figures for its personal inference product. Gemma 4 31B on LiveEquipment Inference measured 192ms, in opposition to Gemini 2.5 Flash at 911ms, GPT-5.5 at 966ms, GPT-4.1 at 1,006ms, and the identical Gemma 4 31B through OpenRouter at 1,876ms.
LiveEquipment is clear in regards to the mechanism, which makes the declare extra credible than most. It runs Gemma behind SGLang with speculative decoding, and intentionally under-packs every GPU so queueing delay stays low. A heat request, it says, begins returning tokens in round 100ms. The tradeoff is value, at $1.20 per 1M output tokens.
The identical put up experiences TTFS throughout full conversations: 354ms for Gemma 4 31B on LiveEquipment, 1,034ms for Gemini 2.5 Flash, 1,088ms for GPT-4.1, 1,267ms for Gemini 3.0 Flash, and 1,404ms for GPT-5.5.
Capability numbers accompany it. On IFBench, independently scored by Artificial Analysis, Gemma 4 31B scores 75.6% in opposition to GPT-5.5 at 75.9%, GPT-4.1 at 43%, and Gemini 2.5 Flash at 39%. On τ²-bench, GPT-5.5 leads at 93.9% with Gemma 4 31B at 76.9%.
Layer 2: Speech-to-Text and Turn Detection
For voice, STT latency just isn’t transcription pace. It is how lengthy after the consumer stops speaking the pipeline is aware of the consumer stopped speaking.
Artificial Analysis measures two issues on its streaming STT leaderboard, each ranging from a SileroVAD-detected finish of speech: time to first partial transcript, and time to last transcript. Its AA-WER Streaming index attracts on roughly 8 hours of audio, weighted AA-AgentDiscuss 50%, VoxPopuli 25%, Earnings-22 25%.
Vendor-published latency figures:
| Model | Claim | Source sort |
|---|---|---|
| Deepgram Flux | ~260ms p50 end-of-turn detection at defaults | Vendor docs |
| Deepgram Nova-3 | Sub-300ms streaming latency | Vendor docs |
| AssemblyAI Universal-Streaming | ~300ms immutable phrase emission | Vendor |
| Cartesia Ink-2 | 100ms transcript latency | Vendor |
| Speechmatics Voice SDK | 0.451 ± 0.022s end-of-speech to finals | Vendor inner instrument |
Deepgram Flux is probably the most architecturally attention-grabbing entry. It folds end-of-turn detection into the popularity mannequin reasonably than bolting a VAD on high. Deepgram states this may reduce agent response latency by 200–600ms versus a standard STT-plus-VAD pipeline. It exposes eot_threshold (0.5–0.9), eager_eot_threshold (0.3–0.9), and an EagerEndOfTurn occasion that permits you to begin the LLM early.
That final functionality issues greater than the uncooked quantity. If you’ll be able to start technology on an keen sign, you progress LLM TTFT off the important path fully when the prediction is correct.
AssemblyAI Universal-Streaming inverts the standard partials-then-finals mannequin by emitting immutable transcripts. AssemblyAI reported 307ms median phrase emission in opposition to 516ms for Deepgram Nova-3 in its personal 2025 measurement. Its docs additionally suggest utilizing unformatted transcripts for voice brokers, since formatting arrives later and hardly ever modifications LLM habits.
Accuracy claims listed here are contested and vendor-published. AssemblyAI experiences Universal-3.5 Pro Realtime at 6.99% WER on the open Pipecat voice-agent benchmark, forward of Google Chirp3 at 9.04%, ElevenLabs Scribe v2 at 9.76%, and Deepgram Flux at 15.58%. Run it your self earlier than treating it as settled.
LiveEquipment additionally paperwork preemptive generation, which begins the LLM on a partial transcript. The caveat is actual: if the reply has to be regenerated after the ultimate transcript, you burn tokens and save nothing.
Layer 3: Text-to-Speech Time to First Audio
This is the place vendor numbers diverge most sharply from what customers expertise.
ElevenLabs states Flash v2.5 delivers roughly 75ms. Its personal docs qualify that fastidiously: 75ms refers to mannequin inference time solely. The firm’s latency concepts page goes additional, itemizing community round-trip at usually 20–200ms relying on geography, and noting that the majority audio gamers buffer earlier than playback, with 500ms buffering being widespread. It additionally states that Eleven v3 just isn’t constructed for real-time, and recommends Flash v2.5, Flash v2, or Multilingual v2 for its Agents Platform.
Cartesia states sub-90ms TTS and 100ms transcript latency for Sonic-3.6 and Ink-2. Marktechpost’s coverage of the Sonic-3.6 release flagged each as vendor-stated mannequin latency, not measured end-to-end spherical journeys. Cartesia beforehand claimed 82ms end-to-end time-to-first-audio for Sonic 3.5. Sonic runs on state house fashions reasonably than transformers, which scale linearly reasonably than quadratically with sequence size.
On high quality, the Artificial Analysis Provider Voice arena is blind-listener Elo, retrieved August 30, 2026:
| Model | Elo | Price per 1M chars |
|---|---|---|
| Cartesia Sonic 3.6 | 1,288 | $49.00 |
| SpeechifyAI Simba 3.2 | 1,243 | $10.00 |
| Alibaba Qwen-Audio-3.0-TTS-Plus | 1,243 | $27.60 |
| Inworld Realtime TTS-2 Flash (preview) | 1,228 | $10.40 |
| BreezeBlue Breeze TTS 2 (open weights) | 1,220 | $34.00 |
| ElevenLabs v3 Conversational | 1,215 | $50.00 |
| Google Gemini 3.1 Flash TTS | 1,210 | $18.30 |
| ElevenLabs Flash v2.5 | 1,083 | $50.00 |
The hole between Sonic 3.6 at 1,288 and Flash v2.5 at 1,083 is the standard value of the low-latency tier that the majority brokers really run on.
Layer 4: Speech-to-Speech Time to First Audio
Speech-to-speech fashions collapse STT, LLM, and TTS into one go. Fewer spherical journeys ought to imply decrease latency.
LiveEquipment is cautious right here, noting that realtime fashions are not guaranteed to be faster in every case, and {that a} well-tuned pipeline might be extremely aggressive.
The knowledge helps that warning. From the Artificial Analysis speech-to-speech leaderboard, TTFA measured on Big Bench Audio, retrieved August 30, 2026:
| Model | TTFA | Speech reasoning | Task success | S2S Index |
|---|---|---|---|---|
| Deepslate Opal | 0.44s | 85% | — | — |
| Gemini 2.5 Flash Native Audio Dialog | 0.63s | 69% | — | — |
| Grok Voice Think Fast 2.0 High | 0.70s | 97% | 94.7% | 79.0% |
| Grok Voice Fast 1.0 | 0.78s | 93% | — | — |
| Qwen3.5 Omni Flash Realtime | 0.79s | 59% | 29.1% | — |
| OpenAI GPT-Realtime-1.5 | 0.81s | 81% | 85.1% | 70.3% |
| OpenAI GPT Realtime Mini (Oct ’25) | 0.81s | 64% | 79.6% | 56.8% |
| OpenAI GPT-Realtime-2.1 Mini Minimal | 0.85s | 63% | 76.7% | 52.8% |
| Google Gemini 3.1 Flash Live Minimal | 0.96s | 71% | 74.6% | 63.9% |
| OpenAI GPT-Realtime-2.1 Minimal | 0.97s | 87% | 89.4% | 70.3% |
| Amazon Nova 2.0 Sonic (Mar 2026) | 1.14s | 88% | 57.1% | — |
| OpenAI GPT-Realtime-2 (High) | 1.14s | 97% | 89.8% | 73.6% |
| OpenAI GPT-Realtime-2.1 High | 1.21s | 96% | 91.5% | 73.9% |
| Google Gemini 3.1 Flash Live High | 2.99s | 97% | 71.8% | 71.5% |
| OpenAI GPT-Realtime-2.1 Mini High | 4.28s | 75% | — | — |
Grok Voice Think Fast 2.0 High is the standout on this board: 0.70s TTFA with 97% speech reasoning and 94.7% process success.
The reasoning-effort penalty is seen inside single mannequin households. Gemini 3.1 Flash Live strikes from 0.96s to 2.99s between Minimal and High. OpenAI’s GPT-Realtime-2.1 strikes from 0.97s to 1.21s, shopping for 2.1 share factors of process success.
OpenAI shipped gpt-realtime-2.1 and gpt-realtime-2.1-mini in early July 2026, and said that improved caching reduce p95 latency by no less than 25% throughout its Realtime voice fashions. Tail latency is what makes a telephone agent really feel damaged, so that may be a extra helpful declare than a median enchancment could be.
The functionality hole
Daily’s benchmark quantifies why most manufacturing brokers nonetheless use cascaded pipelines. On its aiwf_medium_context check, GPT Realtime scored 86.7% in opposition to GPT-4.1 at 94.9%. Ultravox 0.7 was, in Daily’s evaluation, the primary speech-to-speech mannequin to carry out properly on lengthy multi-turn conversations, and it’s open weights.
Artificial Analysis additionally benchmarks 4 vendor “default cascaded methods,” which is helpful context for what the platforms really ship: Deepgram Voice Agent (Nova-3 + GPT-4o Mini + Aura-2), ElevenLabs Agents (Scribe v2 Realtime + Gemini 2.5 Flash + Eleven Flash v2), Cartesia Line (Ink + Gemini 2.5 Flash + Sonic), and Inworld Realtime (Inworld STT 1 + Gemini 2.5 Flash + Inworld TTS 1.5 Mini).
Three of the 4 run Gemini 2.5 Flash. That is a revealing consensus.
Reference Budgets
Assembled from the verified part figures above. These are planning estimates, not measurements of a operating system.
Aggressive cascaded pipeline, US-hosted, colocated:
| Stage | Budget |
|---|---|
| Transport and media (WebRTC) | 50–150ms |
| STT + end-of-turn (Flux at defaults) | ~260ms |
| LLM first chunk (sub-0.5s tier) | 230–500ms |
| Sentence completion at 250+ tok/s | ~100ms |
| TTS first audio + community | 150–300ms |
| Total | ~790ms–1.3s |
That lands at or barely above the 800ms goal, which matches Kwindla’s framing that 800ms is tight however achievable.
Speech-to-speech, single mannequin:
| Stage | Budget |
|---|---|
| Transport and media | 50–150ms |
| Model TTFA (minimal reasoning tier) | 700ms–1.0s |
| Total | ~750ms–1.15s |
Comparable, with much less observability and, per Daily’s benchmark, a measurable functionality hole on instrument calling and instruction following.
What to Do With This
- Pick the metric your structure is bounded by. If a TTS mannequin sits downstream, optimize TTFS, not TTFT. That means TTFT and tokens per second collectively.
- Colocate earlier than you optimize fashions. LiveEquipment charges agent-model colocation as very high impact, above mannequin selection. If you employ SIP, maintain the trunk geographically shut too.
- Cap reasoning effort explicitly. It is the biggest single lever within the tables above, and it’s a configuration flag on most fashionable endpoints.
- Budget for instrument calls. Kwindla notes that any flip with a instrument name roughly doubles LLM latency. LiveEquipment recommends limiting
max_tool_steps, consolidating exterior API calls, and enjoying a pondering sound so silence just isn’t the consumer’s solely suggestions. - Instrument earlier than you tune. The LiveEquipment Agents SDK exposes
e2e_latency, LLM time to first token, and TTS time to first byte per flip. Pipecat exposes the equal by means ofenable_metricsand observers. Store the logs externally and watch for regression. - Measure p95, not simply p50. OpenAI’s personal major enchancment in July 2026 was a tail-latency discount, as a result of that’s the place voice brokers break.
- Watch for infrastructure footguns. LiveEquipment paperwork that self-hosted brokers on AWS burstable occasion sorts reminiscent of t3 or t4g can hit extreme latency and turn-detection timeouts even at apparently low CPU utilization.
Key Takeaways
- Fastest independently measured first chunk on a 10k-token workload: Baseten serving gpt-oss-120b at 0.23s, per Artificial Analysis.
- Throughput and TTFT are totally different merchandise: Cerebras hits 1,697 tok/s however 0.49s TTFT; Inception’s Mercury 2 hits 770 tok/s at 3.07s.
- Vendor latency claims like ElevenLabs’ 75ms and Cartesia’s sub-90ms are mannequin inference time solely, excluding community.
- Reasoning effort is the only largest TTFT lever: Gemini 3.1 Flash Live goes 0.96s to 2.99s between Minimal and High.
- TTFT alone doesn’t predict how an agent feels. Time-to-first-sentence does, as a result of speech synthesis wants a full clause.
Sources
- Artificial Analysis: LLM API Providers Leaderboard
- Artificial Analysis: Performance Benchmarking Methodology
- Artificial Analysis: Speech to Speech Leaderboard
- Artificial Analysis: Streaming Speech to Text Leaderboard
- Artificial Analysis: Text to Speech Provider Voice Leaderboard
- LiveKit: Understand and Improve Voice Agent Latency
- LiveKit: Latency Optimized Inference, Gemma 4
- LiveKit: Voice Agents
- Daily: Benchmarking LLMs for Voice Agent Use Cases
- Daily: Advice on Building Voice AI
- Deepgram: Migrating from Nova-3 to Flux
- Deepgram: Measuring STT Latency
- AssemblyAI: Introducing Universal-Streaming
- ElevenLabs: Understanding Latency
- ElevenLabs: Latency Optimization
- Cartesia: Sonic-3.6 and Ink-2
- OpenAI: Realtime and Audio Guide
- aiewf-eval benchmark source
The put up Lowest-Latency Inference APIs for Voice and Realtime Agents: A Time to First Token TTFT-First Benchmark appeared first on MarkTechPost.
