content

The BFSI Voice AI Evaluation Checklist

A procurement scorecard for NBFCs, banks, insurers, and hospitals in India. Score every vendor against these questions. If a vendor cannot answer them fluently in the first meeting, you are not their target customer and they are not yours.

2026-09-1311 min read

The BFSI Voice AI Evaluation Checklist

A procurement scorecard for NBFCs, banks, insurers, and hospitals in India. Score every vendor against these questions. If a vendor cannot answer them fluently in the first meeting, you are not their target customer and they are not yours.


How to use this

Twenty questions across four sections. Score each answer:

  • 0 — evasive, deflected, or answered with marketing language
  • 1 — partial, vendor-dependent, or requires custom work you'll pay for
  • 2 — architectural, documented, and demonstrable today

Total possible: 40.

  • Under 24 — this is a horizontal API stack with compliance retrofits. Passing an RBI inspection with them will be your project, not theirs.
  • 24 to 32 — a serious contender for a controlled pilot, but expect significant forward-deployed engineering from your side.
  • Over 32 — a vendor built for regulated enterprise. Move to reference calls and a 30-day pilot on a controlled portfolio segment.

Do not accept live demos in place of answers to these questions. Demos are theatre. Production is a different problem — one the industry's best testing platforms fail to catch 13% of the time even on the easiest use case.


Section A — Data residency and compliance architecture

A1. Where does the audio physically travel during a live call, and can you produce a network diagram showing every hop stays inside our regulatory perimeter?

Why it matters. RBI's Outsourcing Direction, DPDP, and IRDAI have made audio residency architectural, not a policy checkbox. If audio transits vendor infrastructure outside India, you have a violation regardless of what the contract says.

Green flag. The vendor produces an unprompted architecture diagram. Every component — STT, LLM inference, TTS, transcript store, embeddings, analytics — is inside your compliance perimeter or a named Indian region (Mumbai / Hyderabad).

Red flag. "We're SOC 2 certified." "Data is encrypted in transit and at rest." "We have DPAs in place." None of these answer the question.

A2. Name the exact database, region, and encryption approach for call recordings, transcripts, embeddings, model checkpoints, backups, and audit logs.

Why it matters. RBI Outsourcing Direction typically requires 5-year recording retention for banks and NBFCs, 3 years for insurance. Vector stores like Pinecone default to AWS us-east-1 unless the customer pays extra for Mumbai. Vendors often have no idea where their own embeddings live.

Green flag. Named database, named region, named encryption approach for every one of the seven storage stages listed above.

Red flag. "It's all in the cloud." "We use industry-standard security." "Let me get back to you on that."

A3. If our DPO receives a right-to-erasure request under DPDP Section 12, walk us through exactly how a specific customer's data is deleted from your system — recordings, transcripts, embeddings, training data, analytics, backups — within the statutory window.

Why it matters. Right-to-erasure is not a checkbox. Vendors who trained on your data often cannot delete it at all — because the model has already learned from it. That is a DPDP violation waiting to be discovered by an auditor.

Green flag. A documented process with named systems and a certifiable proof-of-deletion artifact.

Red flag. Silence, or the phrase "the model has already learned from it."

A4. Are our calls being used to train your foundation model, our tenant-specific model only, or a shared model across your customer base?

Why it matters. If your calls train the vendor's shared model, you are paying to generate training data for a competitor's next quarter of AI improvements. Under DPDP purpose limitation, that is also very likely a violation. "Federated learning," "aggregated insights," and "anonymised improvements" are all training on your data.

Green flag. Our model is our model. Your calls train only your instance. Written into the contract.

Red flag. Any variant of "we improve the platform for everyone using your data."

A5. Can you produce a compliance architecture whitepaper covering RBI Fair Practices Code, DPDP, IRDAI script requirements, and TRAI DND — before we sign?

Why it matters. If the vendor cannot ship this document in week one, your CRO's team will build it from scratch during three months of due diligence. That is a three-month delay you can price into the deal.

Green flag. The whitepaper exists and is shared under NDA immediately.

Red flag. "We'll put something together for you."


Section B — Testing, measurement, and improvement

B1. On what dataset was your evaluator benchmarked against human ground truth, and what is your F1 score on Expected Outcome / First Call Resolution in Indian English and Hindi?

Why it matters. The Andres et al. 2026 benchmark established 0.826 F1 on Expected Outcome as the current industry ceiling — for US English customer support on a leading US voice agent. No vendor serving Indian BFSI has published equivalent numbers on Indian dialects. Ask.

Green flag. A published benchmark with confidence intervals, human-rated, on Indian language conversations.

Red flag. "Our accuracy is over 95%." Accuracy alone is meaningless when 76% of outcomes are naturally positive — a system that says "yes" to everything scores 76%.

B2. Show us your evaluator's Mean Absolute Error on CSAT prediction against 60+ human-rated conversations from a live deployment.

Why it matters. MAE of 0.542 (the current industry-leading number per the same benchmark) means predictions deviate by half a satisfaction point on average. MAE above 1.0 means your automated QA is off by a full satisfaction point — the difference between compliance and complaint.

Green flag. A number under 0.7 on your language and use case, from a real production deployment.

Red flag. "We don't measure it that way."

B3. Show us an improvement curve — F1 score, hallucination rate, resolution rate — month-over-month for an existing production deployment on that customer's calls.

Why it matters. If the model doesn't improve on your data, you are renting last month's generic model forever. Every enterprise conversation is either compounding on your side or on the vendor's — there is no third option.

Green flag. A twelve-month trajectory chart from a real customer, showing metric improvement measurably tied to that customer's call volume.

Red flag. "The foundation model gets better over time." That's their improvement, not yours.

Why it matters. One hallucinated APR in a collections call, one fabricated coverage claim in an insurance renewal, is a regulatory event. Text-based hallucination leaves a paper trail. Voice hallucination does not — the customer heard it, the transcript may or may not have captured it, and the caller cannot scroll back.

Green flag. A hallucination-specific F1 score on regulated scripts, in your language, with adversarial test coverage.

Red flag. "Our LLM has very low hallucination rates."

B5. What is your Word Error Rate on code-switched Indian speech — Hindi to English mid-sentence, Hinglish, Tanglish, Punjabi-English?

Why it matters. Code-switching is standard in Indian retail banking. WER under 8% on code-switched dialogue is the current bar for production-grade deployment. Horizontal STT trained primarily on US English does not clear it.

Green flag. Published per-dialect WER numbers with sample recordings you can listen to.

Red flag. "We support 57 languages." Support is not accuracy.


Section C — Voice identity, specificity, and language

C1. Can we train the TTS on our own brand voice — a specific voice actor, our tonal identity, regional variants — and own the resulting voice model?

Why it matters. In BFSI, the voice is the brand for millions of customers who only ever interact by phone. If every NBFC on the same vendor sounds identical, you have flattened decades of brand identity into a default TTS voice. The CMO will veto this once they hear it — the question is whether they veto before or after you sign.

Green flag. Yes, trained on your voice talent, exclusive to your instance, owned by you.

Red flag. "You can pick from our 40 voices." Or, "we can adjust the pitch."

C2. How does your model learn our specific product vocabulary, regulatory scripts, and escalation patterns without us writing a 40-page prompt?

Why it matters. If the answer is prompt engineering, you have signed up to maintain a Rube Goldberg machine. Real specificity comes from training, not from context injection.

Green flag. A documented fine-tuning or continual-learning pipeline that ingests your call data and improves the model weights.

Red flag. "You'll define scenarios and personas in our workflow builder."

Why it matters. Regulatory updates are frequent. If it takes engineering time to propagate a script change, you have a compliance lag. If you cannot prove the change happened, you have an audit gap. Both surface at inspection.

Green flag. A minutes-to-hours propagation window with an auditable log of when each call started using the new script.

Red flag. "We'll deploy an update in our next release cycle."

C4. What is your accuracy on South Indian English, Marathi, Bengali, and Gujarati specifically — and can we hear sample calls in each?

Why it matters. Language coverage claims are often built on Latin-script transliteration or Google Translate wrappers. The dialects that matter for pan-India BFSI are exactly the ones horizontal vendors underperform on. Sample calls are the disambiguator.

Green flag. Named languages, per-dialect WER, five-minute sample recordings you can listen to on the spot.

Red flag. "We support all major Indian languages."


Section D — Commercial and risk terms

D1. Is pricing per-minute, per-conversation, or outcome-based — and can you tie the price to a CFO-tracked metric like deflection rate, collections recovery, or appointment fill?

Why it matters. Per-minute pricing incentivises the vendor to keep customers on the line. Outcome-based pricing aligns them with your P&L. In regulated collections, per-minute pricing is often the first RBI Fair Practices flag.

Green flag. Outcome pricing tied to a metric your CFO already reports.

Red flag. "It's $0.11 per minute, no minimum."

D2. If we terminate the contract, do we retain the trained model, the training data, and the transcripts — or does it stay with you?

Why it matters. Your calls become a strategic asset over time. If the vendor keeps the model on termination, you have paid for two years of improvements you cannot take with you. This is the operational definition of lock-in.

Green flag. Model, weights, training data, and transcripts are yours on termination, exported in a documented format.

Red flag. "The trained artifacts remain proprietary to the vendor."

D3. Are you comfortable being named in our quarterly RBI outsourcing report as a critical vendor, and will you sign the same undertakings our human DRA agency signs?

Why it matters. RBI's outsourcing circular now treats voice AI vendors as regulated outsourced partners. The compliance officer signs the same form for you as for the human DRA agency. Vendors who won't sign these undertakings are non-starters.

Green flag. Yes, and they've done this at other NBFCs already.

Red flag. "Our legal team will need to review the specific undertakings."

D4. Who is the customer reference in Indian BFSI who has taken you through a full RBI inspection, and can we speak to their CRO?

Why it matters. The single most powerful diligence signal. Vendors who have survived an RBI inspection have already done the compliance work. Vendors who haven't will do it on your project, at your cost.

Green flag. A named reference, named CRO, available for a call within a week.

Red flag. "We have several customers who are close to that stage."


Final scoring

TotalVerdictNext step
Under 24 / 40Horizontal API stack with compliance retrofits.Do not proceed.
24 – 32 / 40Serious contender, expect FDE work from your side.Reference calls + controlled pilot.
Over 32 / 40Built for regulated enterprise.Reference calls + 30-day pilot on 5,000 accounts.

If a vendor scores over 32, the next diligence step is not a demo. It is a 30-day pilot on a controlled portfolio segment (typically 5,000–10,000 accounts in the 0–30 DPD bucket for NBFCs), run in parallel with your existing operation, measuring recovery rate, connect rate, compliance score, and CSAT head-to-head.


Prepared by Telenow. Voice AI bifurcates. Horizontal APIs win pilots. Vertical, on-premise, company-specific voice brains win production. Telenow builds the second one.

Benchmark reference: Andres et al., "Testing the Testers: Human-Driven Quality Assessment of Voice AI Testing Platforms," arXiv:2511.04133v2, January 2026.

Last updated 2026-09-13

Frequently asked questions

How do you evaluate a voice AI vendor for a bank or NBFC?+

Score every vendor across four dimensions — data residency and compliance architecture, testing and measurement rigour, voice identity and language specificity, and commercial and risk terms. Do not rely on live demos, which hide production failures. Use a structured 20-question scorecard with green-flag and red-flag guidance to compare vendors on the same 0–40 scale.

What questions should NBFCs ask voice AI vendors before signing?+

Five essentials: where does audio physically travel; who owns the trained model on termination; what is your F1 score against human ground truth on Indian languages; will you sign the same undertakings our human DRA agency signs under RBI outsourcing rules; and who is your BFSI reference customer that has survived an RBI inspection?

Can voice AI be deployed on-premise for BFSI compliance?+

Yes — and for RBI, DPDP, and IRDAI compliance in India, it usually has to be. On-premise or VPC deployment means audio never leaves the enterprise's regulated perimeter. This is architectural compliance, not policy compliance. API-based voice platforms structurally cannot meet this bar because audio must transit vendor infrastructure to reach their models.

What is a good Word Error Rate for Indian voice AI?+

The bar for production-grade Indian voice AI is under 8% WER on code-switched dialogue — Hindi to English to a regional language mid-sentence. Horizontal STT models trained primarily on US English do not clear this. Ask any vendor for per-dialect WER numbers on South Indian English, Marathi, Bengali, and Gujarati, with sample recordings you can listen to.

How should voice AI vendors be priced for BFSI?+

Outcome-based pricing tied to a CFO-tracked metric — deflection rate, collections recovery, or appointment fill — beats per-minute pricing on two fronts. It aligns vendor incentives with your P&L, and it avoids the RBI Fair Practices flag that per-minute pricing raises in regulated collections. If a vendor only offers per-minute pricing, treat it as a red flag.

How should an NBFC pilot a voice AI vendor?+

The standard pilot: 5,000–10,000 accounts in the 0–30 DPD bucket, Hindi and English at minimum, run in parallel with the existing human collections team for 30 days. Measure recovery rate, connect rate, compliance score, and CSAT head-to-head. Do not expand without this parallel comparison — vendor benchmarks are not your benchmarks.

What happens to a trained voice AI model if we switch vendors?+

This is the lock-in test. Under correct contract terms, the trained model weights, training data, and transcripts stay with the enterprise on termination, exported in a documented format. If the vendor retains "trained artifacts" post-termination, the enterprise has paid for months or years of improvements it cannot take with it. Get export rights in writing before signing.

Does an RBI outsourcing report need to name a voice AI vendor?+

Yes. RBI's outsourcing circular now treats voice AI vendors as critical outsourced partners for regulated financial services. The compliance officer signs the same undertakings for the voice AI vendor as for the human DRA agency. Vendors who will not sign these undertakings — or who have never done so at another regulated entity — are non-starters.

More from Solutions

$0.99 free credit on signup

The BFSI Voice AI Evaluation Checklist

Sign up free and get $0.99 in credit — no card required. Connect your number, pick a template, and go live in minutes.