What an AI voice agent is, in floor terms
An AI voice agent on an outbound floor has one job: hold a short, scripted conversation with a stranger well enough that the stranger stays on long enough to qualify. Everything else, the voice, the dashboard, the integrations, only matters if that conversation works. This post looks at the six places where it either works or does not.
Under the hood there are three parts. Speech recognition turns what the caller says into text. A language model decides what to say back, inside the limits of your script. A voice engine speaks the reply. When all three are quick and well tuned, the caller hears a conversation. When one is slow or wrong, the caller hears a machine.
Many bots sold to Pakistani and Indian floors as AI are still soundboards: recorded clips triggered by keywords. They work until the caller says something the clip list did not expect, and then they loop or go silent. Ask any vendor directly which kind you are buying.
Latency: the pause that gives it away
Latency is the gap between the caller finishing a sentence and the bot starting its reply. People answer each other almost instantly on the phone. Leave a long gap and the caller says "Hello? Hello?" and talks over the reply, and from there the call gets messy.
Latency comes from several places at once: how fast the model thinks, how fast the voice is generated, and the network path between your dialer and the vendor's servers. That last part is why a demo number proves little. Test on your own trunk, from your own floor, at the hour you will actually run, because a call that feels fine at noon can lag at peak.
When you test, do not just judge whether it feels fast. Count how often you find yourself saying "hello?" into a pause, and how often you and the bot talk over each other. Those two counts, taken across twenty test calls, tell you more than any latency figure on a spec sheet.
Interruptions: does it stop when the caller talks?
Real callers interrupt all the time: to answer early, to ask who is calling, to say they are not interested. A bot that finishes its sentence while the caller is talking sounds like a robocall, however good the voice. The feature that handles this is called barge-in, and it is the single best predictor of whether callers stay on.
Good barge-in has two halves. The bot stops speaking the moment the caller starts, and then responds to what was said rather than resuming its line. It should also be tuned to speech, not any sound, so a television in the background does not cut it off mid-word. Test both halves on purpose.
Pay special attention to interruptions that carry a decision. "Not interested" said over the opener, or "Stop calling me" said mid-question, has to be caught even though the bot was talking at the time. A bot that misses those because it was busy speaking will rack up complaints that never show up in a transfer report.
Accents, on both ends of the line
On the caller's side, the US is not one accent. There is the rural South, Boston, New York, older callers with soft voices, people on speakerphone in a car. Speech recognition is strong on a narrow vocabulary (yes and no, numbers, state names, carrier names) and weaker on long, open answers. That suits fronting, where most answers are short. Our live transcript is tuned for that narrow vocabulary, and the recording stays the authoritative record.
On the bot's side, the voice should sound like it belongs on a US call. That is a voice choice, not a training problem, which is one advantage a bot has over a new fronter still working on their accent. Language matters too. Our bots speak English today, with Spanish coming, so Spanish-preferring leads still need a bilingual human for now.
Your own network matters too. If the bot hears the caller through your trunk, poor audio on your side (packet loss, jitter, a congested link at peak) makes the caller harder to understand. That is a network fix, not a bot fix, and it hurts your human agents in exactly the same way.
Scripts and objections: how much freedom should the agent have?
This is the real design choice. More freedom makes the bot sound natural. It also makes it more likely to say something you never approved, like a benefit claim on a Medicare call. The setup that works on regulated campaigns is a fixed spine with flexible joints.
The spine is fixed: the disclosure, the qualifying questions and their order, the knock-out answers and the exit. The joints flex: answering "Who is this?" or "How did you get my number?" from approved facts, handling a short list of allowed rebuttals, and coping with answers given out of order. Anything outside that, and the bot transfers or ends the call politely. Our guide to writing a script a bot can run shows how to draw those lines.
Write the side-question answers down as carefully as the questions themselves. "Who is this?", "Where are you calling from?" and "Is this going to cost me anything?" come up on almost every campaign, and the answer the bot gives should be the one your compliance team would give, word for word where it matters.
The handoff to a human
On a transfer floor the call is only as good as its last ten seconds. The voice agent should dial the closer queue while the caller is still on the line, wait for a human to answer, introduce the caller with what they qualified on, then drop off. If nobody answers, it books a callback instead of leaving the caller in a ringing queue. Closers notice the difference within a shift.
On verification the handoff runs the other way. The closer finishes the sale and passes the caller to the verifier, which reads the compliance statements and confirms the details on a recorded line. What matters there is that the captured data lands in your CRM and dialer comments without anyone retyping it, and that a failed verification is marked failed rather than passed to keep the numbers up.
Platform builders vs done-for-you bots
Voice agent vendors split into two camps. Platforms such as Vapi, Retell, Bland and Synthflow give you the tools to build an agent: you write the prompts, connect the telephony and usually pay by the minute. If you have engineers, or an ops person who enjoys building, that control is worth a lot. Our comparison with Vapi says so plainly.
Done-for-you vendors build the bot from the script your floor already reads, connect it to your dialer and hand it over for approval. You give up some control over the flow and you get time back. Most call centers we talk to have no engineers and do not want to hire any, which is why that is the model B3 Voice sells, including custom bots for flows our standard campaigns do not cover.
A test call that shows you the truth
Before you sign, ring the agent yourself and run this exchange. It takes three minutes and covers most of the failure points above. Write down what the bot does at each step.
Run it at least twice: once on a clean line and once on speakerphone with a television on, at the hour you plan to go live. Send your notes to the vendor afterwards. How quickly they fix what you found tells you what support will be like once you have signed.
- Bot opens. Interrupt it in the first sentence with "Sorry, who is this?" It should stop, answer and carry on.
- Bot asks your first question. Answer two questions at once: "I'm 67 and I'm in Ohio." It should not ask your age again.
- Bot asks the next one. Say "Hmm, hang on," and stay silent for five seconds. It should wait, then check in gently.
- Bot continues. Ask "Is this a robot?" It should answer honestly, in the wording your counsel approved.
- Bot pushes on. Say "I'm not interested." It should use one approved rebuttal at most, then let you go.
- Call again. Say "Take me off your list." It should confirm, end the call and log a DNC disposition, with no rebuttal at all.
- Call a third time and qualify. Listen to the transfer from the closer's side. Were you introduced, or dropped?



