
How fast is an AI voice agent on real calls?
Across 171 real phone calls on our Retell agents, the median AI voice agent took 1.2 seconds to start answering after the caller stopped talking, and the typical p90 turn was 1.6 seconds. That's the honest production number. Not the 500ms you'll see in a benchmark chart, and not the 3 seconds people complain about in forums either. Somewhere in between, with a long tail that does most of the damage.
We pulled every phone call from the past 12 months across the voice agents we run: 366 calls total, 171 of them long enough to have end-to-end latency recorded. Inbound receptionists, outbound follow-up, a Spanish-language real-estate agent, a bilingual assistant. Different models, different voices, different jobs. Here's what the numbers actually say, and what we'd change first if your agent feels slow.
How we measured voice agent latency
Retell returns a latency object on every call through its get-call API. The fields that matter:
- e2e — from when the caller stops talking to when the agent starts talking. This is what the caller feels.
- llm — from issuing the LLM request to the first speakable chunk coming back.
- s2s — the same idea for speech-to-speech models, from request to first audio byte.
- tts — from triggering text-to-speech to the first byte of audio.
- asr — how far transcription lags behind the audio.
Each one comes with p50, p90, p95, p99 and the raw per-turn values. We work from p90, not the average, because callers don't experience an average. They experience the one turn where the agent goes quiet for three seconds while they're giving their address.
A few caveats so nobody over-reads this. It's our calls, not a lab. Some agents had hundreds of calls and some had six. And e2e excludes the network hop from Retell to the caller's phone, so the caller hears a bit more than the number shows.
What the 171 calls showed
The headline numbers:
- Median turn: 1,197ms. Half of all calls had a typical turn faster than this.
- Median p90: 1,641ms. On a typical call, one turn in ten took longer than 1.6 seconds.
- 66 of 171 calls (39%) had a median turn over 1.5 seconds. That's the zone where people start saying "hello?" into the silence.
- 45 calls had a p90 over 2 seconds.
- Only 8 calls had a median turn under 800ms.
So if you've been told sub-second is normal, it isn't, at least not without deliberate tuning. Most of the agents we inherited or built early sat right around 1 to 1.5 seconds, and nobody noticed until they listened to recordings.
Speech-to-speech vs cascading: smaller gap than the hype
This was the result that surprised us. The common pitch is that speech-to-speech models (one model hears audio and speaks audio) crush the classic cascade of transcription, then an LLM, then a separate voice. You'll read "500ms versus 2 to 4 seconds" in more than one comparison post.
On our calls:
- Speech-to-speech, gpt-realtime-mini, 94 calls: median turn 1,063ms, median p90 1,331ms. The model itself returned first audio in about 590ms.
- Cascading, GPT-4.1, 23 calls: median turn 1,139ms, median p90 1,354ms. The LLM returned its first speakable chunk in about 500ms, and the voice added roughly 210ms.
That's a 76ms difference at the median and about 23ms at p90. Real, but nowhere near the gap the marketing implies. The realtime agent was also tuned aggressively (max responsiveness, backchanneling on), while the cascade ran mostly on defaults.
Why so close? Because both architectures still have to wait for the caller to finish. Turn detection, deciding that the person is actually done speaking, eats a chunk of every turn no matter how the response gets generated. Speech-to-speech removes two handoffs. It doesn't remove the wait.
And the cascade buys you a lot for those 76ms: a clean transcript for your CRM, the ability to swap the voice without touching the brain, and much easier debugging. Most production agents on platforms like Retell still run cascading, for exactly those reasons. We pick speech-to-speech only when a casual, chatty feel matters more than control, and that's not most client work.
The model is the biggest latency lever
When we lined up the cascading agents by LLM, the spread was bigger than anything else in the data:
- GPT-4.1 agents: first speakable chunk around 500–630ms median across several agents, end-to-end turns 1.1–1.4 seconds.
- Claude Sonnet 4.5 (Spanish-language real-estate agent, 7 calls): first chunk around 1,400ms median, end-to-end turn about 2.1 seconds.
Swapping the model moved a turn by close to a full second. Nothing else we've ever tuned (voice provider, transcription mode, sensitivity sliders) comes close to that. Text-to-speech, for comparison, was mostly 200–300ms on the cascading agents. Transcription lag was usually under 60ms.
This doesn't mean Claude is bad for voice. That agent ran five tools and a fairly detailed prompt in Spanish, and the reasoning quality on property questions was excellent. It means you're paying for that quality in pause length on every single turn, and you should make that trade on purpose. If you're running Claude elsewhere and watching cost too, our notes on Claude API cost optimization cover the prompt-size side of the same problem.
Tool calls are where the long pauses come from
The median tells you whether an agent is generally snappy. The tail tells you whether people hang up. And the tail, in our data, was mostly tool calls.
That same real-estate agent had one 6-minute call where it searched properties seven times, created a lead, booked a visit and updated a status. Median turn on that call: 3.1 seconds. Worst turn: 7.4 seconds. Every search was a round trip to an external listings API while the caller sat in silence.
On the speech-to-speech receptionist the pattern was the same but milder. Calls that triggered a function had a median p90 of about 1,580ms, versus 1,240ms for calls that didn't. Roughly a third of a second added to the slowest turns, just from one lookup.
What we do about it:
- Pre-fetch on connect. Look up the caller's CRM record by phone number the moment the call starts, before the first word. That lookup shouldn't happen mid-sentence.
- Speak before slow tools. A short "let me check what's open" said before a 2-second search turns dead air into a normal pause. Humans do this too.
- Timeouts on everything. If a calendar or listings API takes longer than a couple of seconds, fail gracefully and offer a callback. A frozen agent is worse than an honest "I'll text you the options."
- Keep webhooks fast. If your tool is an n8n webhook, respond immediately with what the agent needs and push the slow work (CRM writes, notifications) to after the response. We covered the pattern in n8n error handling and monitoring.
The weird tail: agents that slow down mid-call
One more pattern worth knowing about. A 60-second inbound call on the realtime agent started normal, around 1.3 seconds a turn, then the last three turns went 4.5 seconds, 7.8 seconds and 10.6 seconds. No tool calls. Nothing the caller did differently.
We can't tell you exactly what happened inside the model on that call, and we won't pretend to. But it's a good reminder that a per-call median can look fine while individual turns fall off a cliff. If you only chart averages, you'll never see calls like this. Pull the raw per-turn values, flag any turn over 3 seconds, and listen to those recordings first.
What we'd change first on a slow agent
If your agent feels slow, this is the order, based on where the milliseconds actually went in our data:
- Measure it properly. Export e2e p50 and p90 per call, split by agent and by "used a tool or not." Ten minutes of scripting against the API beats a week of guessing.
- Check the model component. If llm or s2s is above about 700ms median, test a faster model first. Biggest single gain available.
- Fix tool calls. Pre-fetch, add spoken filler, add timeouts. This is how you kill the 4-second pauses.
- Trim the prompt. Long system prompts slow the first token on every turn. Move reference material into a knowledge base or tool descriptions.
- Then tune turn-taking. Responsiveness and interruption settings matter, but they're a smaller lever. We go deep on that side in why AI voice agents talk over callers.
- Voice provider last. TTS was rarely more than 300ms on our cascading agents. It's the last place we'd look.
Most teams do this list upside down, auditioning voices for a week while the real problem is a 1.4-second model and a slow calendar webhook.
Is 1.2 seconds good enough?
For most business calls, a 1 to 1.3 second median with a p90 under 1.5 seconds is fine. Callers booking an appointment or asking about a service will tolerate it, especially if the agent sounds natural. Where it breaks down is the tail: the turns over 2 seconds are when people repeat themselves, talk over the agent and lose patience.
Our target for client agents: median under 1.1 seconds, p90 under 1.5 seconds, and zero turns over 4 seconds that aren't preceded by a spoken filler. That's achievable on a cascade with a fast model and disciplined tools. It's what we tune toward when we deploy Retell agents for client lead follow-up, and it's also cheaper: fewer wasted seconds per turn shows up directly in the cost per minute.
If you're still choosing a platform, the latency story is one of the things we compared in Vapi vs Retell for agencies.
Get a free automation audit
If your voice agent is live and callers keep saying "hello?" into the silence, send us access to the call logs. We'll pull the real latency numbers, show you which turns are slow and why, and tell you whether it's the model, the tools or the turn-taking. No pitch deck.
Get a Free Automation Audit — or see how we build and tune voice AI agents for clients.
