New: Voice AI Orchestration Benchmarks — Retell, Vapi, Pipecat, LiveKit & more

How do I choose a voice AI testing platform?

Satvik Dixit
Written byAUG 14, 202610 MIN READ
Satvik DixitinExpert verified
Founding Engineer, CekuraMS, CMU

Has stress-tested 5M+ voice agent minutes at Cekura.

How do I choose a voice AI testing platform?

Why Trust Cekura on Voice AI Evals

  • Built by engineers from Google, Apple, Microsoft. Backed by Y Combinator.
  • 60K+ voice AI calls evaluated daily.
  • Native integration for every major voice AI stack: LiveKit, Pipecat, Vapi, Retell, ElevenLabs, Telnyx.

Choose a voice AI testing platform on measured evidence, not feature lists. Compare simulation realism, evaluator accuracy, layer isolation, consistency

Choose a voice AI testing platform on measured evidence, not feature lists. Compare simulation realism, evaluator accuracy, layer isolation, consistency across repeated runs, and compliance posture. Cekura publishes benchmark data showing pass rates fall by as much as 11.8 points when a scenario must pass three consecutive runs instead of one.

Last updated: August 2026 By Satvik Dixit, Founding Engineer, Cekura

TL;DR

  • A single passing run proves almost nothing. Per Cekura's benchmarks, requiring three consecutive passes drops one platform from 88.1% to 76.3%, an 11.8 point fall, while the strongest falls only 2.3 points.
  • Median latency and tail latency rank platforms differently, so a platform can look fast on P50 and still be the one users wait on.
  • Feature parity is not evidence. No public standard existed for measuring whether a testing platform's own simulations and evaluators are accurate until 2025.
  • Compliance posture and testing quality are separate purchases. A SOC 2 report says nothing about whether a platform finds bugs.
  • Multi-agent systems need assertions on handoffs and tool arguments turn by turn, not a transcript score at the end.

What should you actually compare when choosing a voice AI testing platform?

Comparison starts with the property separating a real evaluation from a demo: does the platform measure the layer that actually failed. A voice agent chains speech recognition, a language model, turn-taking and telephony, and a failure in any link presents identically to the caller. A tool extended from text-based evaluation inherits blind spots at the audio layer, so the criteria below assume one built audio-native.

Per Cekura's benchmarks, one agent ran unchanged across six orchestration platforms with a byte-identical, SHA-verified system prompt and four tool definitions, scored by 59 evaluators across four categories, each scenario run three times. Holding the agent constant is what makes the spread attributable to the platform rather than the prompt.

Apply the same standard to any vendor: one that cannot say what its comparison held fixed is showing you a demo.

What to compareWhy it decides the outcomeWhat the benchmark shows
Consistency across runsIntermittent failures are the ones that reach productionpass^1 ranges 98.9% to 88.1%; pass^3 ranges 96.6% to 76.3%
Tail latency, not medianUsers experience the slow turn, not the average oneFastest P50 is 1.73s; the lowest P95 of 2.95s belongs to a platform whose P50 is 2.34s
Layer isolationTells you what to fix, not just that something brokeAgent and prompt held constant; ASR pinned on 4 of 6 platforms
Baseline honestyTuned configurations flatter a vendorTwo platforms run from providers' standard public templates
Behavioural scoringInterruption handling ranks differently from speedInterruption scores run 4.90 down to 4.63 on a 0 to 5 scale
Evidence of evaluator accuracyAn inaccurate grader is worse than no graderNot answerable from platform marketing, see below

Every figure here is a comparison under one fixed harness, not a production success rate.

How do you read a voice AI testing benchmark?

A benchmark is only as good as what it held constant, so start with two questions about the harness itself.

Who ran it, and was the system identical across arms? Cekura's holds the agent, prompt hash and tool definitions constant, and states two of the six platforms ran from the providers' standard public templates, which makes those a default-path baseline rather than a tuned maximum.

What does the pass criterion require? Retell moves 98.9% to 96.6% between one run and three, a 2.3 point drop. ElevenLabs moves 88.1% to 76.3%, 11.8 points. Both figures derive from the published rates by subtraction, and the gap between those two drops is the honest measure of consistency. A vendor quoting a single-run number is quoting the flattering half.

None of these figures describe production. They describe behaviour under one harness.

What should a benchmark disclose about what it did not control?

The useful disclosure is the one a vendor has no incentive to make, so treat its presence as a quality signal in itself.

What did it fail to hold constant? Cekura publishes this. Speech recognition was pinned to Deepgram nova-3 on four of the six platforms; Retell exposes only a coarse recognition mode and ElevenLabs forces its own Scribe engine, so those two results carry their own transcription stack as well as their orchestration. That bears directly on the latency figures, because the fastest median in the set belongs to one of them.

What conditions did it test? VoiceBench notes that evaluations "focus primarily on automatic speech recognition (ASR) or general knowledge evaluation with clean speeches, neglecting the more intricate, real-world scenarios that involve diverse speaker characteristics, environmental and content factors." Ask which of those three a vendor varied, and treat clean-audio results as a floor rather than an expectation.

What does an enterprise voice AI testing platform need that a smaller deployment does not?

An enterprise voice AI testing platform carries obligations beyond finding bugs: where data lives, who touched it, and what happens to a recording containing personal health information.

Cekura documents its enterprise posture in Securing Conversational AI Observability, covering SOC 2 Type II, tenant isolation, redaction at both the transcript and audio layers, data residency across the US, Europe and India, and deployment inside a customer's own cloud. Cekura records actor, action, source and organisational scope in its audit log.

Compliance and capability are separate purchases, and conflating them is the common procurement error. The HIPAA Security Rule at 45 CFR 164.312 requires audit controls outright, with no addressable implementation specification of the kind encryption at rest has, so two vendors can both claim compliance with materially different protections. Cekura's own page on SOC 2 compliant voice AI testing covers the certification question in full. Neither answers whether a platform tests well.

What does multi-agent voice AI testing require?

Multi-agent voice AI testing has to assert on the transfer itself, because that is where these systems fail. When one agent hands to another, control moves completely and context is either preserved or reset, and a transcript-level score cannot see the difference.

Partner tooling shows the shape of the requirement. LiveKit's agent testing documentation describes validating "Specific messages, tool calls, arguments, and handoffs that you assert on, turn by turn," with sequential event assertions rather than a single end-of-conversation judgment.

The practical criterion follows. A platform that scores a finished transcript will report a failed call without identifying that the handoff dropped a captured account number two turns earlier. Cekura scores each turn against configurable evaluators and covers the scenario construction behind that in The Complete Cekura Scenario Testing Guide.

What does voice AI testing cost to run?

Voice testing costs real money per execution, which shapes how much of it you can afford and therefore how much coverage you actually get.

Vapi's voice testing documentation states plainly that "Each test consumes calling minutes from your account," that "Maximum call duration is limited to 15 minutes per test," and that "Voice tests require more time to execute compared to chat tests." Those constraints are structural rather than vendor-specific: a voice test occupies a telephony leg in real time.

That arithmetic drives the consistency question directly. Running each scenario three times triples the bill, which is why single-run pass rates are quoted more often than three-run rates. Ask a vendor which one their headline number is. Per Cekura's benchmarks the difference between them reaches 11.8 points, so the choice of pass criterion moves the answer more than most feature differences do.

How is testing quality itself measured?

Testing quality is the property nobody was measuring, and it decides whether a platform's pass rate, latency and interruption scores can be trusted at all. A platform that generates unrealistic conversations or grades them inaccurately produces confident, useless numbers.

Cekura answers the question by publishing the harness rather than only the scores. Cekura's published methodology states that every platform was measured with 59 automated evaluators across four categories, three runs per scenario, one agent held byte-identical across all six platforms, and gpt-4.1 pinned at temperature 0.

Cekura also records where that harness stops short. The same methodology discloses that speech recognition was pinned on only four of the six platforms, so the rates describe orchestration behaviour rather than model quality. That disclosure is the part worth demanding. The transferable question for a buyer is simple: ask a vendor how they know their own evaluators are accurate, ask which variables were held constant and which were not, and treat any comparison that will not answer both as marketing.

Where does Cekura fit?

Cekura tests voice and chat agents end to end and publishes the measurement method rather than only the result. Cekura simulates full conversations against a live agent over a real telephony path, scores each turn against configurable evaluators, and monitors the same agents in production so failures return to the test set.

Cekura's predefined metrics cover word error rate (WER) and character error rate (CER) for transcription, time to first token (TTFT) and end-of-turn detection accuracy for latency, and behavioural checks including hallucination, relevancy and response consistency, documented in A Developer's Guide to Voice AI Evaluation Metrics. Cekura also sets out the buyer's evaluation criteria independently of its own product in How to Actually Evaluate Voice AI Testing Platforms.

The strongest single reason to run the comparison yourself is that Cekura's benchmark numbers are reproducible against a stated harness. Ask any vendor for the same: the agent held constant, the prompt verified, the pass criterion named, and the run count stated.

Frequently asked questions

What matters most when choosing a voice AI testing platform?

Consistency across repeated runs, because intermittent failures are the ones that reach production. Per Cekura's benchmarks, the same six platforms score 98.9% to 88.1% on a single run and 96.6% to 76.3% when a scenario must pass three consecutive runs. Every figure is a comparison under one fixed harness, not a production success rate.

Does a SOC 2 report mean a platform tests well?

No. SOC 2 describes how a vendor handles your data, not whether its simulations are realistic or its evaluators accurate. The two are separate purchases. Under the HIPAA Security Rule at 45 CFR 164.312, audit controls are required while encryption at rest is addressable, so two compliant vendors can differ materially in protection.

How do I compare latency between voice testing platforms?

Look at the tail, not the median. Per Cekura's benchmarks the fastest median per-turn latency is 1.73s, but the lowest P95 of 2.95s belongs to a platform whose median is 2.34s. Users experience the slow turn. These figures come from one fixed harness rather than production traffic.

What is different about testing multi-agent voice AI?

The handoff is the failure point, so assertions have to cover transfers, tool calls and their arguments turn by turn rather than scoring a finished transcript. LiveKit's testing documentation describes exactly this pattern, asserting on messages, tool calls, arguments and handoffs sequentially.

Why does running each test three times matter?

Because a non-deterministic system can pass once by chance. Per Cekura's benchmarks the drop from single-run to three-run pass rates is 2.3 points for the strongest platform and 11.8 points for the weakest, which is the clearest available signal of consistency. Both figures derive by subtraction from the two published rates and describe one fixed harness.

Test your voice and chat agents with Cekura

Cekura simulates thousands of conversations before you ship and monitors every call in production — catching broken tool calls, prompt regressions, and instruction-following failures before your users hit them.

Ready to ship voice
agents fast? 

Book a demo