Voice AI Completes 38% of Real Tasks. Then You Add an Indian Accent.
Voice AI Completes 38% of Real Tasks. Then You Add an Indian Accent. Navin, co-founder, Telenow In March 2026, Sierra.ai — one of the best-funded enterprise voice AI companies in the world — published a benchmark called
Voice AI Completes 38% of Real Tasks. Then You Add an Indian Accent.
Navin, co-founder, Telenow
In March 2026, Sierra.ai — one of the best-funded enterprise voice AI companies in the world — published a benchmark called τ-Voice. It measured whether current voice agents can actually complete real customer service tasks: returns, exchanges, cancellations, plan changes, billing inquiries. 278 tasks. Three leading voice AI providers — OpenAI, Google, xAI. Two conditions: clean audio and realistic audio.
Here is the number the industry was avoiding.
Under realistic conditions — background noise, diverse accents, natural interruptions, the kind of audio a customer actually generates on a mobile phone — voice agents completed 26 to 38 percent of tasks. The best text model, running on the same problems, completed 85 percent. Voice agents retained 30–45 percent of text capability.
For the customer service category — the one voice AI has raised billions of dollars on the promise of automating — this is the number that decides the next five years. And it was published by a competitor, not a skeptic.
Why this number is worse than it looks
The paper tested three US English domains: retail, airline, telecom. It did not test Hindi. It did not test Hinglish, code-switching, Marathi-accented English, Bengali, or Tamil. It did not test RBI script requirements, DPDP consent capture, or IRDAI product disclosures. It tested American voice agents against American customer service tasks in American accents.
On that easiest possible version of the problem, the best voice agent completes 38 percent of tasks under realistic conditions.
Now add BFSI. Add regulated collections in Hinglish where the customer switches mid-sentence to Marathi. Add the KYC opening turn where the customer spells their name letter-by-letter over a 2G mobile connection with a fan running in the background. Add the RBI Fair Practices requirement that the agent must disclose its identity within 30 seconds, in a language the customer comprehends, with an audit trail that survives inspection.
The 38 percent number does not stay 38 percent. It gets worse.
Four findings that matter more than the headline
The paper contains four findings that any BFSI CIO or CRO needs to read into their next vendor evaluation.
Accent inequality is now measured, not theorised. Sierra tested seven personas including a Bengali accent, a Sichuan Chinese accent, a Senegalese-French accent, and a Maharashtrian Indian English accent. Under diverse accents, xAI lost 38 percent of its clean-condition capability. Google lost 2 percent. This is not a small difference — it is a chasm. And it means the choice of voice AI vendor for a bank with pan-India customers is now a DPDP fairness question and an RBI Fair Practices question rolled into one. Which side of that gap is your vendor on? If they cannot show you per-accent numbers, they do not know either.
Authentication is the dominant failure mode. From the qualitative error analysis: "Agents fail to transcribe names and emails even when spelled letter-by-letter, blocking all downstream actions." Every BFSI call in India begins with KYC or account identification. If horizontal voice AI cannot reliably transcribe a name spelled out over the phone under noise, the entire regulated use case is gated at turn one. Nothing downstream matters if the first turn fails.
Hallucinated completions. The paper documents agents stating out loud "I've updated your shipping address" without ever calling the underlying tool. In retail this is a bad customer experience. Translate to BFSI: "your EMI has been rescheduled," "your KYC is now complete," "your policy has been renewed," "your loan has been closed." Each of those is a regulatory event with real customer harm and — because voice leaves no scrollback — no easy audit trail. The customer heard it. The database says something different. The dispute takes six months to resolve, and the RBI complaint is already filed.
No vendor masters both task completion and conversational dynamics. OpenAI hits 100 percent responsiveness but 6 percent selectivity — it responds to nearly every backchannel, cough, and stray "hold on" from the customer. xAI leads on task completion but interrupts the customer 84 percent of the time. Google is polite (21 percent interrupt rate) but ignores 31 percent of user turns entirely. Every current vendor is world-class on one dimension and unusable on another. There is no "best" voice AI in the current architecture — only different failure modes distributed across different vendors.
This is architectural, not fixable
The tempting response to any of these findings is: "Newer models will fix this." They will not. Not without changing the architecture.
Every current voice AI stack rents STT, LLM, and TTS from three separate vendors and wires them together with a workflow engine. STT converts speech to text. LLM reads the text and generates a response. TTS speaks the response. The three components share no representation of the call. Tone, hesitation, code-switching, the audible sound of a customer losing patience — all discarded at the first hop. The reasoning layer never sees the audio. The speaking layer never sees the customer.
When failures show up in production, teams reach for the same fixes: bigger prompts, more guardrails, terminology dictionaries, workflow branches. Prompts hit 40 pages. Workflows branch into the thousands. Nobody names what is actually happening — it is compensation for a rented, generic brain that cannot be taught what your specific NBFC's collections vocabulary sounds like when a Marathi-accented customer says it over a bad connection.
Sierra's own conclusion, on the last page of the paper: voice agents will match text models only when they can "sustain fluid conversation while reasoning over multi-step tasks in real time — a constraint text agents, which can think silently for as long as needed, do not face."
That is the shared-representation argument, in a paper published by a leading voice AI vendor. The category is telling itself the answer.
Where regulated BFSI goes from here
The next generation of regulated enterprise voice AI is not going to be a horizontal API stack with better prompts. It is going to be a company-specific voice brain — one integrated stack, trained on the enterprise's own calls, running on the enterprise's own compute, improving weekly on the enterprise's own data.
One stack, because STT, understanding, and TTS have to share representations of the call.
On-premise, because audio never leaves the RBI, DPDP, and IRDAI perimeter.
Continual learning on the customer's own conversations, because the only way to get from 38 percent to 90 percent is training, not prompt engineering.
The API vendors cannot build this. Their unit economics depend on multi-tenancy, and on-premise single-tenant deployment breaks their margin model. The frontier labs will not build this either — every hour spent on RBI residency requirements is an hour not spent shipping their next model. This is not a competitive gap. It is a structural one.
Voice AI is about to bifurcate. Horizontal APIs will keep winning demos and running SMB pilots. Vertical, on-premise, company-specific voice brains will win regulated production.
Sierra just published the number that makes the bifurcation inevitable.
38 percent is not a starting point. It is a ceiling.
We are building what comes next.
— Navin Co-founder, Telenow
Reference: Ray, Dhandhania, Barres, and Narasimhan. "τ-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains." Sierra.ai and Princeton Language and Intelligence, arXiv:2603.13686, March 2026.
Last updated 2026-09-16
Frequently asked questions
What is the τ-Voice benchmark?+
τ-Voice is a March 2026 benchmark from Sierra.ai and Princeton Language and Intelligence that evaluates full-duplex voice AI agents on 278 real customer service tasks across retail, airline, and telecom. It measures task completion using verifiable database state changes and tests OpenAI, Google, and xAI voice models under both clean and realistic audio conditions.
How accurate are voice AI agents in 2026?+
Per the τ-Voice benchmark, leading voice AI agents complete only 31–51% of real customer service tasks under clean conditions and 26–38% under realistic conditions with background noise and diverse accents. In comparison, text-based reasoning models complete 85% of the same tasks. Voice retains only 30–45% of text capability.
Why does voice AI fail on Indian and other non-American accents?+
Under diverse accents, xAI's voice AI loses 38% of its clean-condition capability while Google loses just 2%. Speech-to-text models are trained primarily on American English and struggle with regional accents, code-switching, and Indian English varieties. For BFSI serving pan-India customers, per-accent accuracy testing is now a fairness question, not a nice-to-have.
What is voice AI hallucination in customer service?+
Voice AI hallucination is when an agent confidently states an action has been taken — "I've updated your address," "your policy has been renewed" — without actually calling the underlying tool. The τ-Voice benchmark documents this as a common failure mode. In BFSI, hallucinated completions become regulatory events with no scrollback audit trail.
Can voice AI handle KYC and customer authentication?+
Not reliably. The τ-Voice benchmark identifies authentication as the dominant failure mode: agents fail to transcribe names and emails even when customers spell letter-by-letter, blocking all downstream actions. Since every BFSI call in India begins with KYC or account identification, this failure mode gates the entire regulated use case at turn one.
Which voice AI provider is best — OpenAI, Google, or xAI?+
Each excels on a different dimension but fails on another. OpenAI has the fastest latency (0.9s) and 100% responsiveness but responds to nearly every backchannel and cough (6% selectivity). xAI leads task completion but interrupts customers 84% of the time. Google is most polite but ignores 31% of user turns. No provider masters both task completion and conversational dynamics.
Why does horizontal voice AI fail at task completion?+
Horizontal voice AI stacks rent STT, LLM, and TTS from three separate vendors and wire them together. The three components share no representation of the call, so tone, hesitation, and code-switching are lost at every hop. Teams then compensate with 40-page prompts and thousand-branch workflows — architectural failure hidden as prompt engineering.
Is voice AI ready for regulated BFSI in India?+
Horizontal voice AI stacks are not. The τ-Voice benchmark shows 38% task completion on American English customer service. Add Hindi, Hinglish, RBI script requirements, DPDP consent capture, and mobile audio quality, and the number degrades further. Regulated Indian BFSI requires vertical, on-premise, company-specific voice AI trained on the enterprise's own calls.
More from Solutions
Voice AI Completes 38% of Real Tasks. Then You Add an Indian Accent.
Sign up free and get $0.99 in credit — no card required. Connect your number, pick a template, and go live in minutes.