Voice AI on a clinic phone line is one of the few automation cases where the business argument is easy. Calls are missed, evenings are uncovered, and a large share of the conversations are the same six exchanges repeated all day.
The argument is easy. The evaluation is not. Here is what I ask, in the order I ask it, and I have yet to run through the list without the plan changing.
01
Before the vendor: what am I actually fixing
What is the number this is meant to move? Missed calls, first response time, appointments attended, agent hours on repetitive work. One number, named before any demo.
Do I have that number today, for the last eight weeks? Usually not, and that is the first finding. A clinic that cannot produce its abandonment rate by hour of day cannot evaluate anything against it. Collect for eight weeks. The collection alone often shows the problem is a staffing pattern rather than a technology gap, and then there is nothing to buy.
Is the bottleneck the phone or the diary? If calls are being missed because there are no consultation slots to offer, a faster answering system just delivers disappointment more efficiently.
02
Scope: what may it do, what may it never do
I split every use into three buckets before the first demo.
Fine: outbound reminders and confirmations, rescheduling, directions, opening hours, basic qualification questions on an inbound enquiry, post-visit satisfaction checks.
Only with a human close by: pricing conversations, first clinical enquiries, anything where the caller sounds uncertain.
Never: clinical advice, symptom assessment, medication questions, results, complaints, or any call where the person is distressed. Not because the model cannot generate a fluent answer. Because it can, and a fluent wrong answer in a warm voice is worse than a hold queue.
Write the never list down before the pilot, in the contract if possible. It is much harder to add later, when the containment rate is being celebrated.
03
The questions I put to the vendor
Language and accent coverage for my actual patients, tested on recordings from my own line, not their showreel. Indian English across regions, code switching mid sentence, background noise from a clinic reception. Every vendor demos beautifully on clean audio.
Latency, measured end to end, on my network, at my busiest hour. Anything above roughly a second of silence and callers start talking over the bot, which degrades everything downstream.
Interruption handling. Can the caller cut in? What happens when two people speak at once? This is where natural conversation actually lives.
Escalation. How does a caller reach a human? How many turns does it take? Does the human receive the transcript and the context, or does the patient start again? A transfer that resets the conversation is worse than no bot.
Data. Where does the audio go, where is it stored, for how long, who can access it, what is used for training. Patient voice recordings are personal data under Indian law and under the regimes of any country whose patients you serve. Get the answer in writing.
Failure behaviour. What does the system do when the model is unavailable, when confidence is low, when it does not understand twice in a row? "It apologises and tries again" is the wrong answer. The right answer is that it hands over.
Pricing shape. Per minute, per call, per resolution, per seat. Model it at three times your expected volume and again at a third. Vendors price for the volume they hope you reach.
Who owns the prompts and the call flows. If you build six months of refinement into their platform and then leave, what comes with you?
04
The pilot design
One line, one use, one comparison. I would start with outbound appointment confirmation, because the failure is cheap and the baseline is easy.
Run it for at least eight weeks and make sure that window includes a bad week. Every system looks good in a calm week. What I want to see is the Monday after a holiday when volume triples.
Measure five things against the baseline: containment, transfer rate with reasons, abandonment, appointments actually attended, and complaints. That fourth one is the real measure. Containment is a vendor metric. Attendance is a business one, and I have seen containment rise while attendance fell, which means the bot was successfully ending calls that a human would have saved.
Listen to fifty calls yourself. Not a sample the vendor picks. Fifty, chosen at random, including every transfer. Nothing on a dashboard tells you what fifty calls will.
05
The human veto, written down
One paragraph in the governance page. Something like: this system may confirm, reschedule and inform. It may not advise, assess or decide. Any caller may reach a person by saying so once. Named owner, reviewed quarterly.
That paragraph is what turns a technology decision into a clinical one, which is what it always was.
06
Where I land
I would run the pilot. I am not sceptical about voice AI in a clinic, and the version of this technology available now is genuinely good at the narrow, repetitive, outbound end of the work.
What I would not do is buy it to solve a problem I have not measured, or let it near a distressed caller, or sign before I know where the audio is stored. Those three refusals cost nothing and they are the whole difference between a pilot that teaches you something and a subscription you cannot explain in a year.
Questions people ask
Should a clinic use a voice bot to answer patient calls?
For outbound reminders, confirmations and simple qualification, yes, with a fast route to a human. For inbound clinical enquiries and anything involving distress, symptoms or complaints, no. The cost of a wrong answer is not symmetrical.
What should a voice AI pilot measure?
Containment rate, transfer rate and why, abandonment, appointments actually attended, and complaint volume. Compare all five against a baseline you collected before the bot existed, over a period long enough to include a bad week.
What is the biggest risk with voice AI in healthcare?
A confident wrong answer to a clinical question, delivered in a natural voice that the patient trusts. The second risk is a distressed caller trapped in a menu with no way to reach a person.