Skip to content
LLMs & Agents

How do you build a voice AI agent that survives a real phone call?

A voice AI agent is a streaming loop of telephony, speech-to-text, a language model with tools and text-to-speech; the hard parts are latency, interruptions, turn-taking, testing on recorded calls and handoff to a human. Voice or speech appeared in 8 of 417 AI and ML job postings we collected on 29 September 2026, about 2%: a niche specialism.

Nikhil De Silva · Founder, Square 1 AI6 min read

You build a voice AI agent as a loop of four parts, telephony, speech-to-text, a language model with tools, and text-to-speech, and then you spend most of your time on what a demo hides: latency, interruptions, turn-taking, testing on recorded calls and a clean handoff to a person. The model is the easy part. An agent that works in a quiet room with its builder talking fails quickly on a real line, because callers pause, talk over it, mumble account numbers and ask for a human. Surviving that is the job.

It is also a small niche in the job market for now. Of 417 AI and machine learning postings we collected on 29 September 2026, 8 mentioned voice or speech: about 2%. It is a specialism, not a default skill.

What are the parts of a voice AI agent?

Four stages, running as a stream rather than one after another:

Stage What it does What goes wrong
Telephony Gives the agent a phone number, carries the audio (SIP or a provider's API) and sends call events to your server by webhook Dropped calls, one-way audio, call state lost when a server restarts
Speech-to-text Turns the caller's audio into text as they speak Names, numbers and accents misheard; background noise read as speech
The model and tools Decides what to say and what to do: look up a booking, check a calendar, write a record Long answers, confident wrong actions, no confirmation before a change
Text-to-speech Speaks the reply, starting before the full sentence is written Robotic pacing, mispronounced names, speech that cannot be cut off

The alternative is a speech-to-speech model that takes audio in and gives audio out. It can sound more natural, but it is harder to inspect what the agent "heard" or to swap one weak stage, so many teams start with the chain.

Why is latency the first problem?

Because silence on a phone line reads as a fault. In text chat a pause is fine; on a call, a caller who hears nothing starts talking again, and now both sides are speaking at once.

The rule of thumb is to treat latency as a budget you split across the stages and measure every call against, not a number you hope for. The habits that keep it down:

  • Stream everything. Transcribe while the caller speaks, start the model as soon as the turn ends, and speak the first sentence while the rest is generated.
  • Keep replies short. One or two sentences, then a question. A spoken paragraph is a long wait and a long thing to interrupt.
  • Use a filler honestly. "Let me check that for you" while a slow lookup runs beats silence, as long as it is true.
  • Put servers near the callers. Every distant round trip is delay the caller hears.

Measure from the end of the caller's speech to the first sound of the reply, and study the slow calls, not the average.

How do you handle interruptions and turn-taking?

Barge-in is the caller talking while the agent is still speaking, and real callers do it constantly. The agent has to stop quickly, drop the rest of what it planned to say, and listen. That means text-to-speech you can cancel mid-sentence, and conversation state that records what the caller actually heard, not what the model wrote.

Turn-taking is the opposite problem: knowing when the caller has finished. Jump in too early and you cut off someone finding their card; wait too long and every exchange drags. Voice activity detection plus a silence threshold is the usual start, often tuned with a check on whether the sentence sounds finished. Test with slow speakers and noisy streets.

How do you make an agent complete a task by phone?

Treat speech as unreliable input and confirm anything that changes the world:

  1. Collect details in small steps: name, then date, then time.
  2. Read back what the agent understood, especially numbers, spellings and dates.
  3. Act only after a clear yes, through a tool with a strict schema.
  4. Recover when a tool fails or the caller changes their mind, without restarting the call.
  5. Write a structured record a person can check later.

An agent that completes most scripted calls and hands the rest to a person is worth more than one that attempts everything and quietly books the wrong slot.

How do you test a voice agent before callers find the problems?

With evals on calls, not typed prompts. Two sources:

  • Recorded calls with a known correct outcome, replayed through the whole pipeline. Transcription errors only appear when real audio goes through.
  • Simulated callers: a second model plays a caller with a goal and a temperament (impatient, vague, changes their mind, asks for a manager), speaking into your agent through text-to-speech. You can run a large set on every change.

Score each call on task completion, on the rules (confirmed before acting, never revealed another person's details, handed off when asked) and on how long the caller waited. Put the suite in CI so a change that makes the agent worse fails the build. Our piece on evaluating LLM outputs covers rubrics and judges.

When should the agent hand off to a human?

Always when the caller asks, and whenever the agent is out of its depth: repeated misunderstanding, an upset caller, a request outside its tools, or anything with legal, medical or financial weight. Transfer with a summary so the person does not start from zero. Log every handoff: the reasons are your to-do list.

What are the rules on recording and consent?

They differ by country, sometimes by state, and they change. In general terms: tell callers at the start that they are speaking with an automated agent and whether the call is recorded; get consent where the law requires it, which in some places means every party on the call; keep recordings and transcripts only as long as you need them and protect them as personal data; and treat payment details with the extra care their own rules demand. Take legal advice for where you operate, and build disclosure and consent into the call flow from the first version.

Where should you start this week?

Build the smallest loop that talks: microphone in, speech-to-text, one model call, text-to-speech out. Time the gap between when you stop speaking and when it starts. Then interrupt it and see what breaks. Once that works, put it on a phone number, write five scripted calls with a correct outcome each, and run them after every change. New to tool calls? Start with tool use and function calling.

Where does Square 1 teach this?

The Voice AI Agents Bootcamp is twelve weeks, live on Zoom with one instructor, about 15 hours a week, in six blocks that each end in a deployed project and a gate: a browser voice loop under a stated latency budget, an agent answering a real phone number with consent and handoff, a booking agent, a simulated-caller eval suite in CI, a production agent with monitoring and a cost model defended in a recorded viva, and an employer brief with a hiring sprint. It asks for a built LLM application, Python and webhooks. Voice Agents in a Weekend is a short on-demand course, recorded by an instructor and graded by Nova, the AI tutor, ending in one phone agent that completes a booking. Evaluating AI Systems teaches rubrics, golden sets, calibrated judges and CI suites. All three are taking a waitlist. The AI engineer role page lists the day-to-day, and the free agentic AI skill check takes about three minutes.

Questions people ask

What are the parts of a voice AI agent?

Four stages running as a stream: telephony to carry the call, speech-to-text to transcribe the caller as they speak, a language model with tools to decide what to say and do, and text-to-speech to speak the reply. Some teams replace the chain with a speech-to-speech model, at the cost of being harder to inspect.

How do you keep a voice agent's latency low?

Treat latency as a budget split across the stages and measure every call against it. Stream every stage, keep replies to one or two sentences, use honest fillers during slow lookups and run servers close to callers.

What is barge-in in a voice agent?

Barge-in is the caller speaking while the agent is still talking. The agent must stop speaking quickly, discard the rest of its planned reply and record what the caller actually heard, which needs text-to-speech that can be cancelled mid-sentence.

How do you test a voice AI agent?

Replay recorded calls with known correct outcomes through the whole pipeline, and run simulated callers played by a second model. Score task completion, rule-following and waiting time, and run the suite in CI so a regression fails the build.

Is voice AI a common skill in AI job ads?

Not yet. Of 417 AI and machine learning postings collected on 29 September 2026 from Hacker News, Remotive and Arbeitnow, 8 mentioned voice or speech, about 2%.

Free skill check · about 3 minutes

Where do you stand on Agentic AI?

Five questions, and a skill breakdown the moment you finish: your strengths, the gaps to close, and what to learn next from real curriculum.

Start the Agentic AI skill check

Free, with a student account — the check is the first entry in your record.

Learn this by building it

The programmes that teach what this piece covers, each ending in deployed work graded against a rubric you can read.