Voice Agents Don’t Just Need Models That Talk. They Need Models That Decide.
Telenow.ai X JEV
Voice Agents Don’t Just Need Models That Talk. They Need Models That Decide.
Voice AI has improved dramatically.
Speech-to-text is better. Text-to-speech sounds increasingly human. LLMs can follow long instructions, call tools, access knowledge bases and hold surprisingly natural conversations.
As a result, building a voice agent demo has become much easier.
But putting that agent into production exposes a different problem.
During a real customer conversation, the system isn't only deciding what to say next.
It is constantly making smaller decisions:
- What does the customer actually want?
- Which workflow are we currently in?
- Did the customer answer the question?
- Are they objecting, correcting us or simply interrupting?
- Is the conversation progressing?
- Which tool should run?
- Is the customer becoming frustrated?
- Can we safely perform the requested action?
- Should we continue, change strategy or escalate to a human?
- Did the call ultimately succeed?
These aren't primarily writing problems.
They are decision problems.
And today, we often use a text-generation model to solve them.
We use LLMs for decisions because LLMs are available
Consider a common voice-agent architecture:
Caller
↓
Speech-to-Text
↓
LLM
↓
Tool / Workflow
↓
LLM
↓
Text-to-Speech
The LLM gradually becomes responsible for almost everything.
It understands intent.
It decides which tool to call.
It determines whether the customer answered a question.
It decides which workflow branch to enter.
It generates the response.
It may even evaluate whether its own response was correct.
This works remarkably well for many applications.
But there is an architectural mismatch hiding underneath it.
LLMs fundamentally generate language.
Software frequently doesn't need language.
It needs:
intent = PAYMENT_DELAY
customer_frustrated = true
workflow = PROMISE_TO_PAY
escalate = false
confidence = 0.94
We ask a generative model for these decisions, force its answer into JSON or another schema, validate it, parse it and finally turn it back into application state.
What if decision-making were treated as a separate primitive?
System One models introduce an interesting idea
TypeSafe recently introduced Jev, which it describes as a System One model.
Instead of giving a model a prompt and asking it to generate text, you provide state and typed questions.
The model returns structured decisions.
For example, given the current call transcript, customer information and workflow state, a system could evaluate questions such as:
Choice:
What is the customer's intent?
- payment
- complaint
- cancellation
- information
- human_agent
Score:
How frustrated is the customer?
Noul:
Should this conversation be escalated?
The output isn't prose explaining what the model thinks.
It is structured information that software can immediately use.
That distinction becomes particularly interesting for voice agents.
A voice conversation contains hundreds of micro-decisions
Imagine a collections call.
The customer says:
"My salary hasn't come yet. Give me until Friday."
The agent needs to respond naturally.
That is a language-generation problem.
But underneath that response are several different decisions:
intent = PAYMENT_DELAY
promise_to_pay = TRUE
hardship_signal = MODERATE
customer_disputes_debt = FALSE
escalation_required = FALSE
next_workflow = PROMISE_TO_PAY
Those decisions determine what the system should actually do.
Only after making them does the system need an LLM to figure out the best way to communicate with the customer.
This suggests a different architecture.
CALLER
│
STT
│
▼
┌──────────────────┐
│ Conversation │
│ State │
└────────┬─────────┘
│
┌───────▼───────┐
│ Decision Model │
└───────┬───────┘
│
┌───────────┼───────────┐
▼ ▼ ▼
Workflow Tool Escalate
│
▼
LLM
│
▼
TTS
The LLM remains extremely important.
But it becomes responsible primarily for what LLMs are exceptionally good at:
reasoning and communicating through language.
The decision layer answers a different question:
What is happening, and what should the software do about it?
1. Intent routing without another generative step
One obvious example is intent detection.
Instead of asking an LLM:
Analyze this conversation, identify the user's intent and respond with valid JSON using one of the following categories...
the system can ask a typed choice question.
payment
refund
complaint
reschedule
cancel
information
human
The result can directly determine the workflow.
More importantly, confidence can become part of the architecture.
High confidence
↓
Execute workflow
Medium confidence
↓
Ask clarifying question
Low confidence
↓
Reasoning model / human
The application—not another prompt—determines what happens at each confidence level.
2. A real-time supervisor for every call
Human call centers don't rely only on agents.
They have supervisors.
AI agents will likely need something similar.
After every meaningful conversational turn, a decision layer could continuously evaluate:
Did the customer understand?
Is the agent repeating itself?
Is the conversation progressing?
Is the customer frustrated?
Was the objection resolved?
Is required information missing?
Is the current strategy working?
Should the strategy change?
Should a human take over?
Now imagine these signals feeding directly into the runtime.
If repetition becomes high, change strategy.
If frustration crosses a threshold, shorten responses.
If the customer's question remains unanswered, don't continue the workflow.
If escalation confidence becomes sufficiently high, transfer the call.
The agent effectively gets a supervisor watching every conversation in real time.
3. Workflows become adaptive without becoming uncontrolled
Production voice agents often sit between two uncomfortable extremes.
One is completely hardcoded workflows.
They are predictable, but brittle.
The other is giving an LLM significant autonomy.
It is flexible, but enterprises may have less control over exactly what happens.
A structured decision layer creates an interesting middle ground.
Consider:
Customer
↓
Conversation state
↓
Decision layer
↓
Choose approved workflow
↓
Deterministic business logic
↓
LLM generates natural response
The model can understand messy human behavior while the application still controls the actions available to it.
This becomes especially useful when calls involve payments, CRM updates, appointments, collections, account changes or other consequential actions.
4. Guardrails become executable
Many AI guardrails today are instructions inside prompts.
Never promise a refund.
Never disclose sensitive information.
Never provide information outside policy.
But that means a generative model may be responsible for both generating the response and following the instruction governing that response.
Another architecture is possible.
LLM generates candidate response
↓
Decision layer
↓
Is there a prohibited promise?
Does it expose sensitive information?
Is the claim supported?
Does this action require authorization?
Does policy permit this response?
↓
PASS / RETRY / BLOCK
Now policy becomes something application code can enforce.
That distinction becomes increasingly important as voice agents move from demos into financial services, healthcare, insurance and other regulated environments.
5. Evaluate every production conversation
There is another place where structured decision models could become extremely valuable: observability.
Most voice-agent platforms can already provide transcripts, recordings, latency and summaries.
But operators need answers to much more specific questions.
After every call, imagine evaluating 50 atomic questions:
Was identity verified?
Was intent correctly identified?
Was the customer's problem resolved?
Did the agent repeat itself?
Did the conversation enter a loop?
Did the customer become frustrated?
Was the transfer appropriate?
Did the agent make an unsupported claim?
Did the workflow reach the correct state?
Was the objection handled?
Was the next step established?
Did the customer agree to it?
Did the agent follow policy?
Run that across 100,000 conversations and the result isn't merely another dashboard.
It becomes a map of how the agent actually behaves in production.
For example:
Payment objection 18.4%
Agent repetition 11.2%
Incorrect routing 4.8%
Human escalation 7.1%
Unresolved conversations 13.9%
Workflow failure 2.7%
Then go one level deeper.
Workflow node #17
Completion rate: 71%
Customer confusion: 34%
Agent repetition: 29%
Suddenly the team knows exactly where to investigate.
6. The learning loop becomes possible
This may be the most interesting implication.
Humans learn from conversations.
An experienced salesperson handles their thousandth objection differently from their first.
A support agent learns which explanations confuse customers.
A collections agent learns which approaches work for different situations.
AI agents today don't necessarily improve in the same way.
A conversation happens.
It gets logged.
Someone eventually reviews some calls.
They discover a problem.
A prompt gets changed.
Someone tests it.
The new version gets deployed.
The feedback loop remains surprisingly manual.
Structured evaluation could make the loop much tighter.
Conversation
↓
Evaluate behavior
↓
Detect failure
↓
Classify failure
↓
Cluster similar failures
↓
Identify responsible
prompt / workflow / tool / knowledge
↓
Generate candidate improvement
↓
Replay against historical conversations
↓
Measure improvement
↓
Deploy
↓
Observe again
Generation and evaluation become separate systems.
An LLM can propose the improvement.
A decision model can help measure whether the behavior actually improved.
The agent begins to look less like a static application and more like an operational system that can continuously be measured and improved.
7. We may not need an LLM for every step
There is also an economic implication.
A modern voice pipeline can quietly accumulate many model calls.
STT
↓ intent classification
↓ response generation
↓ tool selection
↓ argument extraction
↓ safety check
↓ post-call evaluation
↓ summarization
Not all of these require open-ended generation.
If structured decision models can reliably handle portions of classification, routing, verification, policy evaluation and conversation-state detection, the expensive generative model can be reserved for the places where generation or deeper reasoning is actually required.
That potentially means:
lower latency, lower cost and more deterministic behavior.
For voice AI, where every few hundred milliseconds matter and model calls happen millions of times, that difference compounds quickly.
The voice-agent stack may split into three layers
We increasingly think production voice AI will separate into three distinct responsibilities.
1. Conversation
Models that understand language, reason and communicate naturally.
What should I say?
2. Decision
Models that continuously interpret state and return structured judgments.
What is happening?
3. Control
Software that determines what the system is actually allowed to do.
What should happen next?
Put together:
VOICE AGENT
┌─────────────────────┐
│ Conversation Layer │
│ LLM │
└──────────┬──────────┘
│
┌──────────▼──────────┐
│ Decision Layer │
│ System One models │
└──────────┬──────────┘
│
┌──────────▼──────────┐
│ Control Layer │
│ Workflows + Code │
└──────────┬──────────┘
│
┌──────────▼──────────┐
│ Business Systems │
│ CRM / Payments / │
│ Scheduling / APIs │
└─────────────────────┘
We don't think this means LLMs become less important.
Quite the opposite.
It means we stop asking them to solve every problem in the system.
What we're exploring at Telenow
Building Telenow and watching voice agents operate in production has changed how we think about the problem.
STT, TTS and LLMs are essential.
But once agents encounter thousands of real customers, a much larger operational problem appears around them:
routing, workflows, integrations, evaluation, observability, guardrails, cost, escalation and continuous improvement.
That is increasingly where production voice AI becomes difficult.
We're interested in whether models such as TypeSafe's Jev can become part of that missing decision layer.
Instead of making one large model responsible for the entire agent, the architecture becomes composable:
LLMs talk.
Decision models judge.
Workflows control.
Observability measures.
The system learns from what happens next.
The first generation of voice AI was largely about making machines capable of having a conversation.
The next challenge may be much harder:
making millions of those conversations measurable, controllable and continuously improvable.
That's the infrastructure we're interested in building at Telenow.
Last updated 2026-09-21
Frequently asked questions
What is Jev?+
Jev is TypeSafe AI's System One model. Instead of generating free-form text, it evaluates structured questions against a given state and returns typed decisions that software can use directly.
How can Jev be used in a voice agent?+
Jev can act as a decision layer around a voice agent. For example, it can evaluate intent, select workflow branches, detect escalation conditions, score conversation states, classify objections, evaluate calls and help determine which action should happen next. The LLM can remain responsible for reasoning and generating natural conversation, while Jev handles structured judgments that application code needs.
Does Jev replace the LLM in a voice agent?+
Not necessarily. They solve different problems. An LLM is useful when the agent needs to understand complex context, reason or generate a natural response. Jev is useful when the application needs a structured decision such as a choice, score or true/false-style judgment. A voice stack could therefore use both: LLM → conversation and reasoning Jev → structured decisions Application code → actions and control
Can Jev replace intent classification with an LLM?+
Potentially. Intent classification is a natural example of a structured decision. Instead of prompting an LLM to generate an intent label in JSON, the application can ask Jev to choose among predefined intents and use the returned result directly in its workflow. The suitability still needs to be tested against the intents, languages, latency requirements and real production conversations of the particular voice agent.
Can Jev be used during a live phone call?+
Architecturally, yes. A voice system can evaluate the current conversation state during a call and use structured results to influence routing, workflows or escalation. Whether it belongs in a particular real-time path depends on measured latency, accuracy, reliability and cost for that deployment.
Can Jev detect when a voice agent is failing?+
It could be used to evaluate specific failure signals rather than asking the broad question "Is this call going badly?" For example: Is the agent repeating itself? Has the customer's question been answered? Is customer frustration increasing? Is the conversation progressing? Has the workflow entered a loop? Should the call be escalated? These signals can then be combined by application logic.
Can Jev be used for voice-agent observability?+
Yes, this is one of the more interesting potential applications. A production platform could evaluate every conversation across many atomic dimensions such as resolution, objection handling, workflow completion, escalation, repetition, compliance and customer frustration. Those structured signals could then be aggregated across thousands of calls to identify recurring failure patterns.
What is the difference between a System One model and an LLM?+
LLMs are primarily generative models: given context, they generate sequences of tokens. TypeSafe describes System One models differently: they evaluate typed questions against state and return structured judgments. For a voice agent, a useful conceptual separation is: LLM: What should I say? System One model: What is happening? Application: What should I do?
More from Solutions
Voice Agents Don’t Just Need Models That Talk. They Need Models That Decide.
Sign up free and get $0.99 in credit — no card required. Connect your number, pick a template, and go live in minutes.