An AI agent that understands text has now become commonplace. An AI agent that answers the phone is a different story. As soon as speech is involved, the bar changes: in a conversation, people do not accept a three-second thinking pause and can tell within a single sentence whether they are dealing with something artificial, and the AI Act naturally also obliges companies to be transparent about this.

Behind every telephone conversation with an AI Voice Agent lie three links that together determine whether the conversation feels natural: speech recognition, speech synthesis, and speed. Anyone who understands how those three work also understands why one voice bot is experienced as a breath of fresh air while another is hung up on within ten seconds.

From sound to text: speech recognition

The first step is speech-to-text, also known as speech recognition. The model converts the caller's sound wave into text, which is then read and understood by the AI agent. This is the link where the most quality is gained or lost, because everything that goes wrong here propagates through the rest of the conversation.

In practice, things rarely go wrong with neat sentences in quiet rooms. It goes wrong with a caller in a car, a heavy accent, background noise on a work floor, and especially with technical jargon. Medication names, policy numbers, company names, abbreviations like WIA or AOV: those are precisely the words a generic model doesn't know and that make all the difference in your service provision.

Good speech recognition is therefore never just a matter of choosing the best model. It is a matter of feeding the model your organization's vocabulary, of recognizing when something has been heard with uncertainty, and of politely asking for clarification instead of proceeding on a wrong assumption. An agent that says “I understand you are calling about an appointment, is that correct?” is better than an agent that silently gives the wrong answer.

From text to voice: speech synthesis

The second link is text-to-speech: the agent's answer is converted into a voice. The technology has advanced so much in recent years that a synthetic voice can hardly be distinguished from a human one. Yet many voice agents still sound like a text-to-speech device, and that is rarely due to the model itself.

The difference lies in prosody: stress, rhythm, pauses, and intonation. A person lowers their voice at the end of a statement and raises it when asking a question, and inserts a short pause before listing items. When that nuance is missing, even the most beautiful voice sounds flat.

In addition, there is the pronunciation of anything that is not an ordinary word. Postcodes, dates of birth, and email addresses must always be double-checked with the end user when in doubt. The voice choice itself is also a substantive decision: the voice conducting an absenteeism meeting with an employee requires a different character than the voice confirming a webshop order.

Why Speed Is Everything

The third link is the least visible and the most underestimated: latency—the delay between the moment the caller stops speaking and the moment the agent begins to respond.

In a normal human conversation, that pause is around two hundred milliseconds. If it exceeds one second, people start to wonder: Has the connection been lost? Did he hear me? Do I need to repeat myself? At one and a half seconds, the caller starts speaking again—sometimes right as the agent answers—and that’s when the end user’s patience quickly runs out.

That delay is the sum of all the steps: recognizing the speech, determining that the caller has finished speaking, retrieving information from your systems, formulating the response, converting it to speech, and transmitting it over the telephone network. Each step takes time, and only the total matters.

That is why good voice agents use streaming: they do not wait until everything is ready, but start processing while the caller is still speaking and begin to speak as soon as the first words of the answer are ready. The length of the response also matters. A four-sentence answer not only takes more speaking time, it also takes more thinking time. In speech, brevity is not a stylistic choice, but a technical requirement.

Practical examples

An employee calls the occupational health service to reschedule their consultation appointment. Despite the traffic noise in the background, the agent understands their name, checks the appointment in the absence management system, and offers two alternative times within half a second. The call lasts fifty seconds instead of an eight-minute queue.

A customer calls outside of business hours to inquire about the status of his request. Instead of a voicemail, he is connected to a voice agent who retrieves the file, explains the current status, and schedules a callback at a time that is convenient for him.

A caller asks a question that is too complex or too sensitive to be handled automatically. The voice agent honestly acknowledges this and transfers the call to a representative, including the context of the conversation up to that point, so the caller doesn’t have to repeat their story.

Voice vs. Chat: What's Really Changing

AspectChatVoice
Acceptable response timea few secondssub-second
Margin of error in data entrylayer, the user typeshigher, depending on accent and environment
Ideal answer lengthA paragraph maytwo to three sentences
Correction by the userIt's simple—he reads it backThat's tough—he only hears it once
Effect of an errorvisible and repairableimmediately noticeable and annoying

You don't hear a good voice agent

The quality of an AI Voice Agent isn’t determined by the most pleasant voice or the smartest language model, but by the weakest link in the chain. Perfect speech recognition paired with a slow response results in an awkward conversation. A beautiful voice that mishears the caller’s name loses trust in a single sentence.

At Conversed.ai, we tailor those three elements to each other and to your organization’s specific needs: the terminology you use, the systems that need to be accessed, and the moments when a conversation should be transferred to a human agent. The result is a conversation that the caller doesn’t even have to think about—and that’s exactly the point.

Curious what that sounds like for your organization? Book a demo and hear for yourself.