New: Cekura Voice AI BenchmarksView results

Voice AI testing for healthcare

Atul Jain
Written byOCT 1, 202611 MIN READ
Atul JaininExpert verified
Founding Engineer, CekuraIIT Kanpur

Has stress-tested 5M+ voice agent minutes at Cekura.

Voice AI testing for healthcare

Why Trust Cekura on Voice AI Evals

  • Built by engineers from Google, Apple, Microsoft. Backed by Y Combinator.
  • 60K+ voice AI calls evaluated daily.
  • Native integration for every major voice AI stack: LiveKit, Pipecat, Vapi, Retell, ElevenLabs, Telnyx.

Voice AI testing for healthcare means placing simulated patient calls against a clinical voice agent and scoring every transcript for workflow order, tool calls, transcription accuracy and PHI handling. Cekura runs those calls with synthetic patient profiles, so a test needs no real PHI, and signs a BAA on Startup and Enterprise plans for monitored calls that carry it.

TL;DR

  • Cekura weights a mistranscribed name or number at 1.0 and a filler at zero toward the weighted error count, because overall word error rate hides the words that matter: Afonja et al. found medical-entity WER 4 to 51% worse than overall WER across pre-trained ASR models on accented clinical English, with one medical-tuned exception.

  • Twin Health runs Cekura's full simulation library before every deployment, and its suite checks that the agent never reveals a stored ZIP code or date of birth to a caller.

  • In Cekura's Medicare workflow study, one byte-identical agent scored between 65.2% and 95.7% workflow pass³ across six voice platforms, a score that counts Expected Outcome and mock tool accuracy and excludes infrastructure failures.

  • Cekura's documentation traces missing Retell transcripts and tool calls to HIPAA settings on the Retell account, so check what your platform's storage mode returns before you trust an empty failure report.

  • Cekura's Startup plan includes a signed BAA and DPA, and its PII redaction removes healthcare numbers, dates of birth and medical professional information from transcripts and audio for the fields you list.

What is voice AI testing for healthcare?

Cekura defines voice AI testing for healthcare as pre-launch and regression testing in which a simulated patient, caregiver or member calls the agent and an evaluator scores whether the call followed the clinical workflow. It differs from general voice QA in three ways: order of operations is a safety rule, identity checks gate every action, and the caller may not be the patient.

Cekura gives each simulated caller a test profile carrying a name, date of birth and phone number that match a mock record in your system, so the agent can verify identity and find an appointment without a real chart. Cekura also generates caregiver scenarios on request, in which a family member answers and works through the flow on the patient's behalf.

Twin Health runs its onboarding agent through this kind of suite. Per Twin Health's case study, Cekura's simulations check that the agent never skips screening questions such as dialysis or pregnancy, that it books lab work and care-team visits in the correct chronological order, and that it never reveals a stored ZIP code or date of birth when a caller asks what is on file. Twin Health runs the full simulation library before every deployment.

Why does overall word error rate understate risk on healthcare calls?

Cekura weights mistranscribed names and numbers at 1.0 and fillers at zero toward the weighted error count, because overall word error rate weights every word equally, while the words that decide the outcome on a patient call are drug names, dates and member numbers. A transcript can look accurate and still carry the wrong medication.

Tejumade Afonja of the CISPA Helmholtz Center, Tobi Olatunji of Intron Health and co-authors measured this in Performant ASR Models for Medical Entities in Accented Speech. Across the pre-trained model families they tested, medical WER ran 4 to 51% worse than overall WER, with one medical-tuned cloud model as the only exception. Whisper-large had the best overall WER yet recalled 42% of medication entities and 33% of protected health information entities. The study used clinical English across 93 African accents, so the figures belong to that dataset.

In simulated runs, Cekura's Transcription Accuracy metric scores only the simulated caller's turns, because Cekura generates that speech and holds the exact words spoken, and it counts a negation flip at 1.0 and a verb at 0.5. For what a vendor's accuracy number does and does not prove, see medical voice recognition software.

What should a healthcare voice agent test suite cover?

Cekura builds a healthcare voice agent test suite around the clinical workflow, the backend tools and the privacy boundary, with each case defined by a caller situation and the outcome that passes it.

Test areaExample caller situationWhat passesHow Cekura runs it
Identity verificationCaller gives a date of birth one digit offAgent asks again, stops after two failed attempts, and never reads the stored value backTest profile with mismatched data, plus red teaming
Caller authority"I am calling for my dad"Agent establishes who is speaking and whose consent appliesCaregiver scenario from extra instructions
Clinical orderCaller asks to book before screeningAgent completes required screening questions firstExpected Outcome evaluator
Tool callsCaller asks for an appointment next weekAgent fetches open slots and reads them back correctlyTool Call Success and mock tool checks
Volunteered PHICaller starts reading out a diagnosisAgent does not solicit data it does not needEvaluator on the call transcript
Names and numbersCaller names a medication and doseTranscript keeps the drug name and number intactTranscription Accuracy
EscalationCaller reports chest pain mid-bookingAgent stops scheduling and follows the escalation path your clinical team definedExpected Outcome evaluator
Agent handoffScreening agent hands the member to the lab-booking agentThe next agent receives the verified record and does not re-ask completed questionsExpected Outcome evaluator across the handoff

Cekura's Medicare workflow study shows how far the platform alone moves a result. The same Medicare insurance agent ran on six platforms against 23 evaluators, three repetitions each, and workflow pass³ ranged from 65.2% to 95.7%. Workflow pass³ counts Expected Outcome and mock tool accuracy only. Under strict scoring, which adds infrastructure reliability, one platform fell from 69.6% to 8.7%.

Cekura checks identity rows against the test profile itself. In Cekura's documented example, the expected outcome states that the agent must compare the caller's answer with {{test_profile.dob}}, ask again on a mismatch, and after two failed attempts stop and send the caller to the facility. Generate Scenarios then produces fake-verification callers at scale.

For performance testing, run the same patient workflows as concurrent calls and compare latency and tool-call success before launch. Cekura's pricing page lists 10 concurrent calls on pay as you go, 50 on Startup and custom concurrency on Enterprise.

How do you run HIPAA compliant voice agent testing?

Cekura runs HIPAA compliant voice agent testing by keeping protected health information out of every place it is not needed: synthetic test data on the way in, and stored transcripts that nobody covered by a BAA should read on the way out.

Start with synthetic data. Cekura's test profiles carry invented names and dates of birth that match mock records, so a pre-launch suite processes no real PHI. Production calls do carry PHI. Cekura signs a BAA for that work, and Cekura's PII redaction strips fields such as healthcare_number, dob and medical_professional from the stored transcript and the audio. Redaction does not reach metadata, dynamic_variables or customer_number, so Cekura publishes a client-side script that scrubs those before upload and fails closed if detection breaks.

Then check your voice platform's storage mode, because it decides what a test tool can read. Retell's Basic Attributes Only setting stores no transcripts, recordings or logs, so a later get call request returns none, and Cekura's documentation traces missing Retell transcripts and tool calls to HIPAA compliance settings enabled on the Retell account. Vapi's HIPAA mode applies to the whole organization, restricts it to compliant model, voice and transcriber providers, and limits access to call logs and transcriptions.

Which tools handle voice AI testing for healthcare, and should you build or buy?

Cekura and the other ways to handle voice AI testing for healthcare differ on four points a healthcare buyer decides on: whether tests need real PHI, whether the vendor signs a BAA, whether it scores clinical workflows end to end, and what each test run costs.

OptionNeeds real PHI to testBAAClinical workflow coverageSetup timeCost basis
CekuraNo, synthetic test profilesSigned on Startup, custom on EnterpriseScenarios, tool calls, red teaming, transcription scoringEvaluators generated from your agent's context$0.25 per voice testing minute, or $500 a month on Startup
Staff placing manual test callsDepends on the records testers useNot applicable, no vendor involvedWhatever the tester remembersNoneStaff hours per release
In-house test harnessDepends on your fixturesGoverned by your own HIPAA policiesWhatever you buildSet by your engineering backlogEngineering time plus telephony

Building in-house means owning a caller that dials the agent, an LLM that plays a patient, evaluators per metric, storage for audio a reviewer can open, and a BAA review for every model provider the harness touches. Cekura replaces that harness, and its pricing page lists the signed BAA and DPA with 90-day log retention on the Startup plan, against 30 days on pay as you go.

"With one click, we simulated our 30+ service workflows - scanning every node and edge - saving days of manual QA."

Vichar Shroff, Co-founder and CPO, Confido Health

Per Confido Health's case study, Confido also ran thousands of simulated calls through Cekura before migrating its calling infrastructure, comparing latency, workflow accuracy and tool-call success between the old and new stack.

Frequently asked questions

What are the best tools for voice AI testing for healthcare?

The best tools for voice AI testing for healthcare test with synthetic patient data, sign a BAA, and score the clinical workflow and the transcript together. Cekura does all three: it runs simulated patient calls from test profiles, signs a BAA on its Startup and Enterprise plans, and scores tool calls and workflow order with evaluators and transcripts with a weighted Transcription Accuracy metric.

How much does voice AI testing for healthcare cost?

Cekura charges $0.25 per voice testing minute on pay as you go. The Startup plan costs $500 a month, includes roughly 2,000 voice testing minutes, and adds the signed BAA and DPA. As an illustration, 30 workflows run 3 times at an assumed 3 minutes a call is 30 × 3 × 3 = 270 voice minutes, or $67.50 at the per-minute rate.

Does Cekura handle voice AI testing for healthcare?

Yes. Cekura automates end-to-end voice AI testing for healthcare: it generates evaluators from your agent's context, tests with simulated patients and caregivers, verifies identity checks against test profiles, red teams for data leaks, and redacts the healthcare numbers and dates of birth you list from monitored calls. Twin Health and Confido Health both run their patient-facing agents through Cekura. Cekura tests agent behaviour and does not certify an agent or a programme as HIPAA compliant.

Do you need real patient data to test a healthcare voice agent?

No. A pre-launch suite runs on synthetic patients: Cekura's test profiles carry an invented name, date of birth and phone number that match a mock record in your scheduling or EHR sandbox, so identity verification and lookups behave as they would on a real call. Real PHI enters only when you monitor production calls, which is where a BAA and redaction apply.

How do you fit voice AI testing for healthcare into CI?

Run the full simulation suite as a gate before every deployment, the way Twin Health does, so a prompt change in one agent cannot break a screening question or a handoff elsewhere. Set each Cekura scenario to run more than once, the way Cekura's Medicare workflow study scores a scenario as passed only when all three repetitions succeed, because a clinical workflow that passes one run in three is not safe to ship.

Test your voice and chat agents with Cekura

Cekura simulates thousands of conversations before you ship and monitors every call in production — catching broken tool calls, prompt regressions, and instruction-following failures before your users hit them.

Ready to ship voice
agents fast? 

Book a demo