AI voice-agent evaluation: my work with Acolad
My AI voice-agent evaluation work for Acolad in aviation and healthcare: hands-on testing, conversation analysis and practical recommendations for improvement.
My AI voice-agent evaluation work for Acolad involved hands-on testing and analysis of agents for aviation and healthcare applications. I evaluated the caller experience and reported findings and practical recommendations to the team. The purpose of this kind of review is to turn a conversation with an agent into useful direction for improving it.
An AI voice agent represents a business at a moment when someone wants something done. The caller may need information, a change to an existing request or help reaching the right person. How the agent handles that conversation is part of the product.
That makes evaluation valuable in its own right. A team needs to understand how an interaction lands with the person on the other end, and which changes deserve attention next.
My role in Acolad’s AI voice-agent evaluation
For Acolad, my contribution covered testing, analysis and feedback. I interacted with the AI voice agents, assessed the conversations and returned recommendations about what should be changed.
| Part of the work | My contribution |
|---|---|
| Hands-on testing | Experienced the agents through phone conversations |
| Conversation analysis | Evaluated the interactions from the caller’s perspective |
| Findings and recommendations | Reported observations and suggested changes to Acolad’s team |
The engagement sits alongside the other AI and automation projects in my client work. It represents a different contribution to an AI project: reviewing an existing agent and helping its team decide how to refine it.
The discussion below sets out how I think about that work. The sample situations and feedback are illustrative evaluation examples. They give a business owner something practical to use when commissioning a review of their own agent.
What a useful evaluation should give the team
A useful review connects three things: what happened in the conversation, what that meant for the caller and what the team could change.
The observation needs enough detail to be recognisable. “The conversation was confusing” leaves the builder guessing. A precise observation identifies the question, response or transition that caused confusion. It gives the team a place to investigate.
The explanation should connect that moment to the agent’s job. Did the caller know what to do next? Could they correct information? Was the response relevant to the request? A finding becomes more useful when the business consequence is clear.
The recommendation should describe a better behaviour. It might suggest clarifying an answer, changing the order of questions or improving the way the agent acknowledges a correction. The builder can then choose the appropriate implementation and test it.
This is the level of feedback I would ask for when buying an evaluation: specific enough to guide a change, with a clear reason for making it.
Evaluate the experience of being on the phone
Reading a conversation and hearing it are different ways to review an agent. In an audio interaction, the caller experiences the timing, the pauses and the way the agent responds when they speak. Those details deserve their own attention.
Retell’s testing overview distinguishes text and simulation testing from browser audio and phone testing. Its guidance uses audio calls to examine voice, latency and interruptions, and phone calls to check the telephony path. These methods are useful reference points when scoping an evaluation.
For a business review, I would start with a simple question: does the caller understand what is happening throughout the interaction?
A long answer can leave someone unsure which part to respond to. A question containing several requests can make it difficult to supply the right information. An unexplained pause can cause someone to start speaking again. These are useful scenarios to examine, because they turn an abstract idea such as “conversation quality” into something a reviewer can describe.
The review should also recognise what works well. If a particular question helps the caller answer clearly, the team should know to preserve it during revisions.
Aviation and healthcare change the questions you ask
An industry label is a starting point for evaluation. The more useful detail is the exact task the agent is supposed to handle and the information it is allowed to use.
For an illustrative aviation support scenario, imagine a caller discussing a travel date. They give one date, then correct themselves. A review could examine whether the agent acknowledges the correction and uses the revised date in the rest of the conversation. It could also check whether the caller understands if they have requested a change or actually completed one.
For an illustrative healthcare administration scenario, imagine someone asking about an appointment and then asking to speak to a staff member. The review could examine how clearly the agent explains the next step and whether it follows the agreed route for that request. This example concerns administrative communication.
Neither example requires a reviewer to invent an operational policy. The organisation should supply the intended behaviour, approved information and relevant boundaries before testing begins.
That is what makes an evaluation specific to a business. The standard comes from the job the agent has been given, and the review checks the interaction against it.
Turn observations into practical recommendations
The quality of the feedback matters as much as the act of testing. A team should be able to read a finding and understand what needs investigation without reconstructing the entire conversation.
Here is an illustrative feedback entry for a fictional agent:
Situation: The caller corrected the requested date during the conversation.
Observation: The agent acknowledged the correction, but used the original date in its closing summary.
Caller impact: The caller could leave with the wrong understanding of which date had been recorded.
Recommended behaviour: Use the corrected date in subsequent responses and confirm it before closing.
Follow-up check: Repeat the scenario with different dates and verify that the closing summary consistently reflects the correction.
This format separates an observed behaviour from a possible technical cause. The issue could involve instructions, conversation state or the connection to another system. The reviewer can describe the experience accurately without guessing which component caused it.
It also gives the team a useful way to discuss priority. A confusing word choice and an incorrect closing summary may require different levels of attention. Explain the consequence of each finding so the team can make that decision deliberately.
Keep suggestions proportionate. A recommendation to ask one clearer question may be more useful than asking the builder to redesign an entire conversation. Good feedback reduces uncertainty about the next change.
Give the next version a clear test
A recommendation becomes easier to assess when it includes a way to check the revised behaviour. Decide what you expect the agent to do, then revisit the scenario after the change.
LiveKit’s evaluation documentation distinguishes tests of individual behaviours from evaluation of complete conversations. It covers checking messages, tool use and handoffs as well as interactions across multiple turns. These are complementary ways to organise a review.
For a simple business example, a revised closing message should be checked on its own and as the end of a conversation. The wording may be clear in isolation but confusing after the caller has changed their request. The surrounding interaction gives the message its meaning.
I would also keep successful scenarios in the review set. After changing how the agent handles a correction, run the straightforward version of the same request again. The team then has evidence about both the revised behaviour and the ordinary path it still needs to support.
My AI voice-agent testing checklist turns those principles into a practical sequence, including a sample test record and launch review questions.
What to ask for when commissioning a review
Start by telling the reviewer what the agent is meant to accomplish, who calls it and which tasks need particular care. Share the intended conversation path and the approved next steps. That context helps the reviewer judge whether a response is appropriate.
Agree on the output as well. I would ask for clear observations, the likely effect on the caller, recommended changes and a way to check each revision. If a finding depends on something the reviewer cannot inspect, such as whether a downstream record was actually saved, it should remain an open verification item.
Ask how the review will distinguish a single observed issue from a recurring pattern. Both can be useful, but they support different conclusions. A clear account of the evidence makes the recommendations easier to trust and easier to prioritise.
For a broader view of the service itself, see how an AI phone receptionist fits into a business. If you already have an agent, bring the task it handles and the part of the experience you want to improve to a free 30-minute call. We can identify a useful scope for testing, refinement or a closer look at the workflow.
Questions people ask
What AI voice-agent work did Ivar Knutsen do for Acolad?
Ivar conducted hands-on testing and analysis of Acolad’s AI voice agents for aviation and healthcare applications, then reported his findings and practical recommendations to Acolad’s team.
What does AI voice-agent evaluation involve?
It involves examining how an agent handles a conversation against an agreed purpose, identifying behaviour that needs attention and turning observations into recommendations the team can act on.
What makes voice-agent feedback useful to a development team?
Useful feedback identifies the situation, describes the observed behaviour, explains its effect on the caller and recommends a change with a clear way to check the revised interaction.
Why evaluate voice agents for aviation and healthcare applications?
The evaluation should check how clearly the agent handles its specific task, uses approved information and explains the next step. Scenarios need to reflect the organisation’s intended workflow and callers.
Can voice-agent evaluation help a business outside aviation or healthcare?
Yes. The same questions about understanding, clarity and the next step apply to many business calls. The scenarios and acceptance criteria should reflect the specific task the agent is meant to handle.
Sources
- 1.Retell AI: testing overview · Reference for the distinction between text, simulation, browser audio and phone testing.
- 2.LiveKit: testing and evaluation · Reference for checking individual behaviours and evaluating complete conversations.

Written by Ivar André Knutsen
I build and run AI systems, internal tools and workflow automation. You work directly with me from the first conversation through implementation and support. About Ivar
Want this looked at in your business?
We look at where you want the business to go, what is slowing you down and where AI could make a useful difference. You get a clear recommendation: a tool to try, a focused automation, a broader system or a closer look at the process. Any build is scoped and quoted before work starts. Free, no obligation.
Single automations are quoted on the call: a setup fee plus a monthly retainer to run them. Full systems start at $4,500, fixed scope, fixed price.
Book a free 30-minute call