How to launch a voice AI agent: 9 lessons from real calls
What we learned launching voice AI agents on live phone lines: a dictionary from real calls, a script, code checks, speed, tests and operator handoff.
We run voice agents on live phone lines in two kinds of business. In one, the agent takes phone orders from regular buyers. In the other, it books clients for a medical centre's services. Both speak Ukrainian, write the result straight into the client's system and hand over to a person whatever they cannot handle themselves.
In voice bot ads everything looks simple: connect a model, write a prompt, the agent talks. In practice a working version is made of dozens of small decisions, and we found each of them on real calls. Here are the main ones.
In short
- Customers name products differently from the price list. Without a dictionary built from real calls the agent misses a third of the items.
- Free-form conversation drifts. An agent that follows an agreed script behaves predictably.
- The model talks, but code checks the phone number, the questionnaire and the order.
- One similar-sounding word can hang up the call mid-conversation.
- The agent takes answers from the company knowledge base, not from the model's memory.
- We check every change with a test suite, without live calls.
1. A one-evening demo and an agent on a live line are different things
The first version that answers by voice appears quickly. It is nice to show, but you cannot hand it the phone yet. It interrupts the caller, mixes up items, gets lost when the person asks something off-script, and goes silent when the server answers with a delay.
Everything below is the work between "the agent talks" and "the agent can be trusted with the line". That is where most of the project time goes.
2. Customers name products their own way
On the phone a buyer names a product differently from how it is written in the price list. They shorten it, reorder the words, use colloquial names, speak surzhyk or Russian, and sometimes simply mispronounce it.
We ran almost a thousand product names from a month and a half of real calls through the agent's search:
| Result | Share of names |
|---|---|
| Found right away | about two thirds |
| Found after one clarification | about a fifth more |
| Not found | roughly one in ten |
To close the gap we collected the words buyers actually use for products from real calls. Every doubtful word was approved by the business owner, not by us. The dictionary lives in the database, and the client edits it in their own system without a developer.
Takeaway: before launch, give your contractor recordings of your calls. Customers will not understand an agent trained only on the price list.
3. Free-form conversation drifts, a script holds
One version of the order-taking agent worked in free mode: the model decided on its own when to ask and what to clarify. In tests it handled every call differently, did not acknowledge the caller and could not tell where the dictation ended.
We moved to a script agreed with the client point by point:
- greeting;
- the caller dictates without interruptions, the agent only gives short acknowledgements;
- clarifications only after the word "that's all", one at a time;
- different product groups in separate blocks;
- at the end the agent reads the list back and asks "Is everything correct, shall I record it?".
Small things we also had to spell out: if the caller repeats the quantity the agent suggested, that is a confirmation, not a new item. If the caller asks about availability mid-dictation, the agent answers first and then returns to the order. If the caller is silent for 7 seconds, the agent prompts.
4. The agent talks, code checks the data
The costliest voice agent mistake is a lost phone number or an incomplete request. So we do not trust the model with key data:
- The phone number is assembled by automation from what the person dictated. The agent used to interrupt mid-number; now it does not.
- The pre-booking questionnaire is controlled by code. Without a complete questionnaire the request does not go further.
- A question is not an answer. If a client asks something of their own mid-questionnaire, for example whether the procedure is allowed for them, it does not close the questionnaire item.
- The preferred time is recorded only after the client says "yes".
- A call goes to an operator only after the agent has taken the name and number. Even if the line is busy, the contact is not lost.
5. One word can hang up the call
The agent hangs up when it hears a goodbye. In tests it turned out that some Ukrainian words begin with the same letters as an informal goodbye, and some surnames and ordinary words sound like other service commands. The agent said goodbye mid-conversation.
Now keywords trigger only as whole words. We checked the rule on almost 4,000 real phrases, and none produced a false goodbye.
Same category: the voicemail detector in the voice platform took a live person for voicemail and hung up right after the greeting. You only see this on real calls, so in the first days after launch we listen to recordings every day.
6. Take answers from a knowledge base, not the prompt
When all company knowledge sits in the prompt, the agent gets slower, more expensive and starts making things up. We shortened the prompts and moved prices, terms, addresses and opening hours into the company knowledge base. The agent looks for answers there.
On several hundred real client questions the agent finds the right answer in the base in about 85% of cases. It does not invent the rest: it takes the contact and passes it to an operator. After answering about price, the agent offers to book on its own.
7. Speed: where exactly it slows down
A 4-second pause on the phone feels like forever. We measured every link and saw that the live conversation model thinks 4 to 6 seconds per answer, while our part (knowledge base search, checks, recording) takes 1 to 2.5 seconds.
What helped:
- the agent asks simple questionnaire items in batches, so the answer comes in 2 seconds instead of 4;
- long knowledge moved out of the prompt into the base;
- if the server does not respond, the agent retries once and then says a fallback phrase. Silence on the line is worse than any answer.
8. Test without calls and on real phrases
Calling the agent by hand after every change is slow, and it is easy to miss a regression. We built simulated calls: a synthesized voice plays the "client", and these calls do not go into the call log. The agent records a fast multi-item dictation in full.
Reference answers have a permanent test suite. We run it in full before every change. That is how we catch cases where a fix in one place breaks answers in another.
To make the agent talk like the best manager, we label real conversations of the client's team: tens of thousands of lines from hundreds of calls. They become the reference behaviour.
9. The agent stops for the simplest reasons
One day the telephony started rejecting the agent's calls. The cause was not in the code: the telephony account had run out of money. Another time the model provider balance ran out in the middle of testing.
A voice agent depends on several services at once: telephony, voice platform, model, database. Set up balance and error alerts in each of them, otherwise you will learn about an outage from your customers.
Where the call result goes
A conversation by itself gives nothing; what matters is what remains after it.
- Order taking: the order goes straight into the accounting system with the right prices for this buyer. The manager sees the call recording in the call log.
- Medical centre: the request arrives in the administrators' Telegram group and reads at a glance: who the client is, which service, what needs clarifying. The same "brain" also answers in the website chat.
Do you need a voice agent
| Good fit | Poor fit |
|---|---|
| Many similar calls: orders, bookings, confirmations | Every call is a separate consultation |
| After-hours calls get missed | A few calls a day |
| There is a system to write the result into (CRM, accounting, spreadsheet) | Nowhere to record the result |
| There are call recordings to train the agent on | No recordings and no script |
Where a voice agent pays off, we covered in Voice AI agent: 5 use cases where it pays off.
How we launch a voice agent
- We listen to your call recordings and write down the typical scenarios.
- We agree the conversation script point by point.
- We build a dictionary from the real names and phrases your customers use.
- We run the agent through tests without calls, then give your team a test number.
- We connect it to your line and review calls daily for the first weeks.
FAQ
Does the agent understand surzhyk and Russian?
Yes. That is what the dictionary from real calls is for: the agent records the item correctly however it is named.
Can a call be passed to a person?
Yes. The agent first takes the name and number, then transfers to an operator or creates a request if the line is busy.
Where can I see what the agent said?
Every call is stored with a recording and transcript, and the result goes into your system or Telegram.
How long does launch take?
The first version appears quickly. Fine-tuning on live calls takes a few weeks, and we do not skip that stage.
Let's look at your calls
Would you like to know whether a voice agent makes sense for your calls? Book a call: we will go through your scenarios and show a demo. The first call and demo are free. More about AI agents and pricing on the AI implementation page.
Looking to automate your business the same way? We will show you how on a free intro call.