Home/ Articles/ The 200-Millisecond Number Hiding Inside Enterprise Voice AI Partnerships
Voice Ai

The 200-Millisecond Number Hiding Inside Enterprise Voice AI Partnerships

The gap is what gives it away. Not the timbre, not the breath between phrases, not the sibilance that a good model now renders without that old wet-plastic edge. The gap.

A photorealistic overhead shot of an audio engineer's mixing desk at night, hands nowhere…

The gap is what gives it away. Not the timbre, not the breath between phrases, not the sibilance that a good model now renders without that old wet-plastic edge. The gap. You ask the thing a question, there is a beat before it answers, the beat runs a hair too long, and some pre-verbal part of your brain has already filed the voice under machine before you have processed a single word of the answer.

That beat is what is actually being sold in the current run of enterprise voice AI partnerships — deals where a systems integrator bolts itself to a speech-model vendor and tells its clients it can put a talking agent in front of their customers. The one on my desk as I write this is DXC Technology and ElevenLabs: an IT services firm with decades of enterprise plumbing on one side, a speech synthesis company with a very good vocoder on the other. The announcement language will be about co-innovation and shared go-to-market. The engineering is about a beat.

I want to build this whole piece on one number, because it is the only number here I can source, and because this industry quotes it constantly without ever saying where it came from.

The only number in this piece

About 200 milliseconds.

That is the typical gap between one person finishing a turn in conversation and the next person starting theirs. It comes from a 2009 paper in PNAS by Tanya Stivers and colleagues — "Universals and cultural variation in turn-taking in conversation" — which sampled recorded, naturally occurring talk in ten languages from unrelated families, including Japanese, Korean, Danish, Italian, Dutch, English, Lao, Tzeltal, Yélî Dnye and ǂĀkhoe Haiǁom. Across all ten, the most common response offset clustered in a narrow band around a fifth of a second. Languages differed from each other — the spread between the quickest conversational culture in the sample and the most deliberate one was measured in a few hundred milliseconds, not in seconds — but nobody anywhere was leaving a second and a half of air.

For scale, in terms I use daily: 200 ms is roughly a dotted eighth at 120 BPM. It is longer than a slap-back you would print on a vocal and shorter than the pre-delay on a hall reverb. It sits right at the boundary where a drummer stops hearing tight and starts hearing late.

The reason that number has become the unofficial spec sheet for voice agents is that it is the only empirical anchor anyone has for what "natural" means. Every latency claim you read in a voice AI announcement is being measured, implicitly, against Stivers — usually by people who have never read the paper.

So it is worth being precise about what that study measured, and — this is where both the investment case and the procurement case live — what it did not.

What the study actually measured

The researchers coded question-and-answer sequences in ordinary conversation. Video recordings, informal settings, people who already knew each other. They measured the offset between the end of a question and the start of the response, allowing for negative offsets where the answer began before the question had finished landing.

Two findings matter for anyone signing something.

The first is the convergence. Ten unrelated languages, ten different conversational cultures, wildly different grammars — and the shape of the distribution held. That is what makes the number feel like a floor rather than a fashion. Turn timing behaves less like etiquette and more like a motor skill, closer to how a bassist locks to a kick drum than to how anyone was raised.

The second finding is stranger and more useful. Laboratory estimates of how long it takes a human to plan and launch even a single word run to roughly half a second, and longer for a full phrase. If preparing speech costs 600-odd milliseconds and the observed gap is 200, the arithmetic does not close — unless listeners are predicting where your sentence is going and building their answer while you are still producing it. That is the standing account in the literature; Levinson and Torreira set it out clearly in 2015. Turn-taking is a prediction problem, not a reaction problem.

Stay with that a moment, because it reframes the engineering completely. Humans are not fast responders. Humans are early starters. A system that waits for silence, however quickly it then replies, is playing a fundamentally different game from the one that produced the number it is chasing.

One more caveat, since the paper is being used as a benchmark and was never built as one: these were question-response pairs, a deliberately constrained slice of conversation chosen because the boundary is easy to identify. Real talk is full of turns that are harder to score. The 200 ms figure is a robust central tendency in a specific structure, not a universal metronome for all human speech.

What 200 milliseconds does not measure

Here is the part that falls out when the number crosses from linguistics into a slide deck.

It did not measure customer service. The corpus is casual conversation between people with an existing relationship. A billing dispute at eight in the evening with someone whose flight has been cancelled is a different speech genre with different tolerances. Institutional talk has its own rhythms, and some of those rhythms include pauses that would be rude at a dinner table and are perfectly ordinary when somebody is looking something up.

It did not measure human-to-machine interaction at all. Not one participant knew they were talking to software, because none of them were. Whether people transfer their conversational timing norms onto a machine is a separate empirical question, and the evidence there is thinner and messier than the confidence of the marketing implies.

It did not measure satisfaction. It measured timing. There is no line in that paper connecting response latency to whether anyone felt heard, got their problem solved, or would willingly use the channel again.

And it did not measure resolution, which is the number a procurement team should actually care about. Did the caller get the thing they phoned about without a human being involved? That has nothing to do with milliseconds and everything to do with whether the agent can reach the system of record and write to it.

A close-up photograph of a professional recording studio microphone in a darkened vocal booth…

There is also the inverse failure, which the number says nothing about. Agents that respond too eagerly cut people off mid-thought, and in the testing I have sat in on, testers forgive slowness far more readily than they forgive being interrupted. A pure latency target optimises toward the failure mode humans hate most.

So when an announcement leans on the naturalness of the voice, it is answering the question the study asked and skipping the question the business is asking. Those two may correlate. They may not. Nobody has shown me the work that closes the gap, and I have gone looking.

Why an integrator is quoting a linguistics paper

Strip the adjectives off these deals and the trade underneath is simple and, on paper, sensible.

The model vendor owns the hard audio: prosody that survives a long paragraph, phoneme coverage across languages, streaming synthesis with a short time to first audio, a voice that does not collapse into mush on an unusual proper noun. What the model vendor does not own is the depressing and valuable middle of an enterprise — the mainframe with a schema designed in 1997, the telephony estate, the contact-centre platform three acquisitions deep, the change-control board, the regulator who wants to know which country the audio is stored in.

The integrator owns exactly that, plus the client relationships and the security clearances, and it is under real commercial pressure. The public story at firms of DXC's type for several years has been an older book of business shrinking while the company hunts for work carrying a better margin. I am not going to quote you a revenue figure or a share price here, partly because this piece will still be readable in eighteen months and partly because you should pull the current filing yourself rather than trust a number I typed on a Tuesday. But you can check the shape of it in about ten minutes, and you should, because that shape is the entire context for why a deal like this gets announced with quite so many adjectives.

The honest negative on each side, since a review without one is a brochure:

  • For the model vendor: distribution through an integrator is slow, lower-margin, and puts your audio quality behind somebody else's implementation. If a deployment sounds bad because the telephony leg was built carelessly, the customer blames the voice, not the wiring.
  • For the integrator: you do not own the model. You are reselling someone else's differentiator into accounts where your own margin comes from people and hours. If your client can buy the same voice directly next year at a better rate, what exactly did you build?

That second question is the investment case compressed into one line. Announcements do not answer it. Signed contracts might.

Where the milliseconds actually go

If you want to test any latency claim in this category, read this section twice, because the budget in a real deployment looks nothing like the budget in a demo.

A single conversational turn passes through, roughly: voice activity detection deciding you have stopped speaking; speech recognition producing a final transcript; a language model deciding what to say; often one or more tool calls out to a CRM, an order system or a billing platform; speech synthesis producing its first chunk of audio; and then the network, including the jitter buffer, which exists specifically to trade latency for not sounding like a scratched CD.

Two of those stages tend to dominate, and neither one is what gets marketed.

The first is endpointing. The system has to decide that your silence is a finished turn rather than a breath, a hesitation, or you reading a reference number off the back of a card. Set the threshold tight and the agent talks over people, which is the failure everybody remembers. Set it comfortable and you have spent more than the entire 200 ms budget before recognition has even started. Teams routinely land somewhere in the several-hundred-millisecond range and then discover the figure they promised the client is arithmetically unreachable. This is the single most common place a voice pilot quietly loses its headline latency, and I have never once seen it named in a partnership announcement.

The second is the tool call. The moment your agent has to ask a thirty-year-old system whether an order shipped, your response time is that system's response time plus everything else. The voice model is not the bottleneck. The bottleneck is a SOAP endpoint in a data centre nobody has visited since 2011.

Which is, to be fair, precisely the argument for involving an integrator at all. The people who can make that endpoint answer faster are not the people who trained the vocoder. If these deals work, that is the mechanism by which they work — not because the voice is warmer than last year's.

Barge-in is worth its own line in any evaluation. When a caller interrupts, does the agent stop cleanly, and does it retain what it had already said so the conversation makes sense afterwards? Handling that well is unglamorous state management, and it does more for perceived quality than shaving another eighty milliseconds off synthesis.

The 8 kHz problem nobody puts in the press release

Now the part that makes me, specifically, wince.

Modern speech models render beautifully at high sample rates. You can hear breath, mouth noise, the small transient at the front of a plosive. On monitors it is genuinely impressive, and the demo is usually played on monitors.

Then it goes down a phone line. Traditional telephony is narrowband: the audio is resampled to 8 kHz and squeezed through a codec designed in an era when the goal was intelligibility per bit, not realism. Everything above roughly 3.4 kHz disappears. That is where sibilance lives, where the air in a breath lives, where most of the cues you were paying a premium for live. The voice you signed off in the conference room, played through good speakers at full bandwidth, is not the voice your customer hears.

If the deployment runs in-app or over a WebRTC channel with a wideband codec, you keep far more of it. If it runs over the public phone network, budget for a large share of the perceptual quality you evaluated to be discarded in transit. This is not a criticism of any particular vendor — it is physics and inherited infrastructure. But it means the demo is not the product, and an evaluation conducted on a laptop speaker is measuring the wrong artefact.

A photograph of two people seated across a small table in a bright glass-walled…

Ask for a recording of the agent over the real channel, on the real carrier, at the real codec, with a real background-noise caller on a motorway. If nobody can produce one, that tells you exactly how far along the deployment is.

How I'd decide

These are the criteria I would use whether I were putting money into the stock or a customer-service line into production. The middle column is the question; the right-hand column is what a weak answer sounds like.

Criterion What to ask Weak answer
Latency, measured honestly Caller-perceived, end to end, over the real channel, including endpointing and tool calls, reported at p95 rather than median "Time to first audio is under X" — a component, not a turn
Containment How is a resolved contact defined, and who signs off on that definition Deflection rate: calls that left, not problems that ended
Voice rights Who owns a cloned or custom voice, what consent record exists, what happens to it when the contract ends "That's covered in the platform terms"
Indemnity Who carries liability if a synthetic voice is misused or a claim arrives Silence, or a cap smaller than the deployment
Data and recordings Where audio and transcripts are stored, retention period, whether either trains anything "Enterprise-grade"
Cost over twelve months Per-minute or per-character rates modelled at projected volume, plus integration hours, plus year two A unit price with no volume model attached
Exit Can you take prompts, transcripts and evaluation sets to another vendor An export that loses the voice and every hour of tuning

The pattern across all seven: the differentiator is almost never the audio. It is contract terms and integration depth. Which happens to be exactly what an integrator should be good at, and exactly what never shows up in a demo.

The watching brief

If you hold the stock, or you are trying to judge whether the AI story at any services firm is real, ignore announcement volume — voice AI partnerships cost very little to announce. Watch four things instead, in rough order of how much each would move me.

Named production references with call volumes. Not a pilot, not a logo slide. A client willing to say publicly that a meaningful share of its inbound contact goes to an agent, with a number attached and a date on it.

Whether it appears in bookings rather than press. Signed contract value in the relevant segment, disclosed as such. If this work becomes material, it eventually surfaces in segment reporting. If it never surfaces, it was marketing.

Renewal after the first term. Voice pilots either renew or die quietly. Second-year renewal is the honest quality metric, because it is the one customers vote on with budget rather than enthusiasm.

Whether the firm builds any intellectual property of its own. Reselling a partner's model is revenue protected by someone else's moat. Orchestration, evaluation tooling, domain-tuned agents that survive a model swap — that is a position that holds. If the answer two years from now is still "we implement somebody else's speech models," the multiple should say so.

None of that requires an opinion about speech synthesis. It requires waiting, which most people find expensive.

Who this is for, who should skip it

Enterprise procurement: worth a serious pilot if you have high call volume, a narrow band of repetitive intents, and — the actual prerequisite — clean API access to the systems holding the answers. If your systems of record are hostile, the voice will be excellent and the deployment will still fail.

Investors: treat a partnership as a signal about intent, not about earnings. The interesting question is whether the firm converts integration skill into repeatable, priced work before the model vendors go direct into the same accounts.

Skip it if your call mix is mostly complex, emotionally loaded, or regulated in a way that requires a human on the line. Skip it too if the business case rests on headcount reduction alone. The deployments I have watched hold up were aimed at coverage — overnight, overflow, the queue at two in the morning — rather than at replacing the day shift.

And if you are a creator reading this because you use these same voice models for narration or game dialogue: watch the enterprise licensing terms closely. Terms drafted for a bank's contact centre have a habit of migrating down into consumer tiers, and the direction of travel is usually toward more restriction on what a synthetic voice may be used for commercially, not less.

The unsettled part

I keep circling back to the prediction finding. Humans hit 200 milliseconds not by responding quickly but by starting early — modelling where your sentence is heading and launching before you land it. Almost no production voice agent does this. They wait for silence, then react. The entire category is chasing a number produced by a mechanism it has not built.

Maybe that does not matter. Maybe a fast reactive system is perceptually indistinguishable from a predictive one, and the final hundred milliseconds is engineering vanity dressed up as user experience. Maybe people extend machines a completely different timing budget — more patience, or considerably less — and the human baseline was never the right target in the first place.

The research does not settle this. There is real work on how people time their speech with artificial partners, and it points in more than one direction depending on the task, the stakes, and whether the person knows what they are talking to. Nobody has published the thing that would actually decide it: a large, boring, real-world study connecting turn latency to resolution rate and repeat contact in a live contact centre with real customers who did not consent to being interesting.

Until someone does, every latency claim in this category is an aesthetic argument wearing a lab coat.

So here is the question I would put to any vendor, any integrator, and honestly to myself before forming a view on the stock: if the agent solves the caller's problem in ninety seconds but pauses half a beat too long on every turn, has it failed — and by whose measure?

Not sure which tool to use?

Compare the top AI music and sound tools side by side — honest reviews, real pricing, no sponsorships.

Compare the Tools
N

Nova Reyes

Editor, The Signal

Nova Reyes edits The Signal and reviews AI music tools after a decade scoring indie games and short films; still owns four broken synthesizers. More by Nova Reyes →