New: Voice AI Orchestration Benchmarks — Retell, Vapi, Pipecat, LiveKit & more

Voice AI testing pricing comparison

Shashij Gupta
Written bySEP 18, 20269 MIN READ
Shashij GuptainExpert verified
Co-founder & CTO, Cekura

Has stress-tested 5M+ voice agent minutes at Cekura.

Voice AI testing pricing comparison

Why Trust Cekura on Voice AI Evals

  • Built by engineers from Google, Apple, Microsoft. Backed by Y Combinator.
  • 60K+ voice AI calls evaluated daily.
  • Native integration for every major voice AI stack: LiveKit, Pipecat, Vapi, Retell, ElevenLabs, Telnyx.

A voice AI testing pricing comparison turns on five drivers: testing minutes, repeat runs per scenario, metric evaluations, concurrent call capacity and seats. Cekura publishes its rates: $0.25 per voice testing minute, a headline $0.05 per monitored call, and on pay as you go one seat free then $30 per month each, with ten seats included on Startup.

TL;DR

  • Cekura bills voice testing at $0.25 per minute and monitored calls at a headline $0.05 each, metered as 0.2 credits per metric scored, with one seat free and $30 per month for each seat after it on pay as you go, and ten seats included on Startup, rates read on 17 September 2026.
  • Cekura's Startup plan includes 10,000 credits for $500, which is $0.05 per credit, the same effective unit price as pay as you go, so the plan buys concurrency, seats and retention rather than a cheaper minute.
  • Repeat runs, not scenario count, are the largest controllable line in a testing budget, because one run does not settle whether a scenario passes.
  • Cekura scores metric evaluations separately from call minutes, at 0.2 credits per metric run, so a wide metric set raises the cost of every call in the suite.
  • Cekura bills overages, when they are enabled, at twice the committed credit rate for the tier and caps concurrency by plan, at 10 calls on pay as you go and 50 on Startup, so volume and concurrency are pricing decisions before they are infrastructure ones.

What actually drives the cost of a voice AI testing platform?

A voice AI testing bill is a usage bill with five inputs, and scenario count is rarely the largest of them. The inputs are minutes of simulated conversation, the number of times each scenario is repeated, the number of metrics scored on each call, the concurrency the plan allows, and the number of seats. Only the first is obvious, which is why quoted comparisons that stop at a headline per-minute rate understate the real figure. Cekura meters minutes and metric evaluations as separate lines rather than folding them into one number, and repeat runs multiply both, so an estimate can be built before a contract is signed. Work the multiplication through on your own suite: 80 scenarios at two minutes each is 160 minutes on a single pass, and scoring those same calls against ten metrics adds 800 metric evaluations on top of the call time. Cekura's guidance on voice AI agent cost and performance optimization works the same trade from the agent side.

How do voice AI testing platforms bill, and what does each model cost you?

Voice AI testing bills take three shapes: metered usage, a committed bundle and a per-seat charge. Each shape moves risk between buyer and vendor. Metered usage suits a team with uneven volume, while a committed bundle suits a steady nightly suite and buys a cheaper unit only where the vendor actually discounts it. Cekura combines all three and publishes each on its pricing page rather than quoting on request. A generic calculator cannot run this comparison for you, because the inputs are your suite's, so price each line against your own scenario count and repeat policy.

Billing lineWhat it metersWhat raises itCekura's published rate, read 17 Sep 2026
Voice testingMinutes of simulated conversationLonger scenarios, more repeat runs$0.25 per minute, metered as 5 credits per minute
Chat testingMessages sent by the testing agentLonger conversation paths$0.025 per reply, metered as 0.5 credits per message
Production monitoringMetrics scored on each call imported from production, plus transcription only when a recording arrives without a transcriptMore metrics per call, growth in live call volumeHeadline $0.05 per monitored call, metered as 0.2 credits per metric run plus 0.1 credits per minute of transcription
Metric evaluation on test callsEach metric scored on a simulated call, charged on top of that call's minutesAdding metrics to the suite0.2 credits per metric run
SeatsUsers with platform accessAdding users beyond the plan's included seatsPay as you go 1 free then $30 per month each; Startup 10 included; Enterprise custom
Concurrent callsSimultaneous test calls the plan allowsLoad tests, parallel nightly suites10 on pay as you go, 50 on Startup, custom on Enterprise; extra phone lines $10 per line per month, prorated daily
OverageCredit usage past the monthly allocationUnplanned volumeTwice the committed credit rate for the tier, when overages are enabled

Cekura's documentation sets out every credit rate and works both examples through. Cekura prices concurrency twice, as a plan limit and again as the minutes every parallel call bills, and its load testing guide shows how fast that compounds: a frequency of N runs each selected evaluator N times in one cycle, so ten evaluators at frequency five place fifty concurrent calls.

Why does an honest pricing comparison have to budget for repeat runs?

Repeat runs are the multiplier that turns a small suite into a real bill, and they are not discretionary. One run does not settle whether a scenario passes. Cekura's frozen benchmark study ran the same 82 caller scenarios three times against each of 8 configurations, retaining 246 calls per configuration and 1,968 across the cohort, and its leading configuration passed all three runs on 62 of those 82 scenarios. That is roughly one scenario in four failing at least one of three attempts, on the strongest performer in the cohort, with calls that never connected counted as failures. Published research points the same way. A 2025 study re-evaluating eight models over three independent runs reports that single-run leaderboards are brittle, with 10 of 12 slices inverting at least one pairwise rank against the three-run majority, and it adopts two runs as a cost-effective default with a third run optional for high-precision reporting. Price your suite at two to three times its single-pass minutes.

What does Cekura charge, and what is not included in the rate?

Cekura charges $0.25 per voice testing minute, $0.025 per reply for chat testing and a headline $0.05 per monitored call, rates identical on pay as you go and on the $500 per month Startup plan. Volume discounts appear only at the Enterprise tier. The Startup plan's 10,000 included credits cover roughly 2,000 testing minutes, which works out at $0.05 per credit, the same effective unit price as pay as you go, so the $500 buys 50 concurrent calls, ten seats, 90-day log retention and a signed BAA rather than a cheaper minute. On pay as you go Cekura bills seats separately from usage, one free and $30 per month for each seat after it, and its documentation is explicit that adding a team member does not automatically increase the credit allocation. Two costs sit outside the rate: the runtime your own agent platform bills while a test call is live, and the engineering time to maintain scenarios. Both belong in any platform choice.

When is building your own voice agent test harness cheaper than buying one?

Building your own harness is cheaper only while the suite stays small and the team already owns the telephony path. The cost that catches teams out is not the first harness, it is the second year: scenario upkeep after every prompt change, a scoring layer that has to stay calibrated against human judgement, and infrastructure that can place calls concurrently. Published work gives a sense of the scale. The Holistic Agent Leaderboard reports 21,730 agent rollouts across 9 models and 9 benchmarks at a total cost of about $40,000, and that figure is compute alone, carrying no engineer time and no telephony. The honest version of the trade is that buying converts an unpredictable engineering commitment into a metered line, and it costs a margin on every minute to do so. That upkeep is what Cekura's regression testing path automates.

Frequently asked questions

Which voice AI testing platform should I choose on price alone?

None of them, because the headline per-minute rate is one input of five, and the repeat count multiplies it two or three times before metric evaluations are added. Compare vendors on how they meter repeat runs, metric evaluations and concurrency, and on whether the rates are published at all. Cekura publishes its full rate card, which lets you model a suite before talking to sales rather than after.

What should a voice AI testing suite actually cost per month?

Work it from minutes, not from a plan name. Multiply your scenario count by mean call length, then by two or three repeat runs, then by the per-minute rate, and add metric evaluations separately. At Cekura's $0.25 per minute, an 80-scenario suite of two-minute calls run three times nightly is 480 minutes a night before metrics, which is $120 a night and $3,600 across a 30-day month at the pay as you go rate, seven times the Startup plan's included 2,000 minutes.

How much does voice AI testing add to a CI pipeline's cost?

It adds the minutes of whatever subset you gate merges on, multiplied by your repeat count and your merge frequency. Teams usually gate on a smaller suite than they run nightly for this reason. Cekura meters a CI run identically to any other run, so the cost is predictable once the gating suite is fixed.

What do enterprise compliance and audit requirements add to the price?

Cekura prices compliance by tier rather than as a line item. A signed BAA and DPA arrive with the $500 Startup plan, while SSO, SCIM, audit logs, data residency and VPC or on-premise hosting are Enterprise features priced on a custom contract. Longer log retention, which audits usually need, also scales with tier.

Does Cekura handle voice AI testing pricing comparison?

Cekura does not compare other vendors' pricing for you. Cekura publishes its own rates in full, meters voice minutes, chat messages and metric runs as separate lines, and reports credits consumed by each metric in the billing dashboard, so the figures you need on one side of a comparison are checkable rather than quoted. Load testing is priced on the same metered minutes.

Test your voice and chat agents with Cekura

Cekura simulates thousands of conversations before you ship and monitors every call in production — catching broken tool calls, prompt regressions, and instruction-following failures before your users hit them.

Ready to ship voice
agents fast? 

Book a demo