New: Cekura Voice AI BenchmarksView results

Drift detection for voice AI agents

Dileep Chagam
Written byOCT 1, 202610 MIN READ
Dileep ChagaminExpert verified
Founding Engineer, CekuraIIT BombayEx-Apple

Has stress-tested 5M+ voice agent minutes at Cekura.

Drift detection for voice AI agents

Why Trust Cekura on Voice AI Evals

  • Built by engineers from Google, Apple, Microsoft. Backed by Y Combinator.
  • 60K+ voice AI calls evaluated daily.
  • Native integration for every major voice AI stack: LiveKit, Pipecat, Vapi, Retell, ElevenLabs, Telnyx.

Drift detection for voice AI agents is the practice of scoring every production call on fixed metrics and alerting when a metric moves away from its recent baseline. Cekura does this with trend alerts that compare each new call against an exponentially weighted moving average, and with failure modes tracked per agent version.

TL;DR

  • A voice agent drifts when the same agent starts producing different outcomes because the LLM, the speech stack, the callers or your own prompt changed underneath it, and Cekura detects it by scoring call outcomes rather than model internals.

  • Cekura scores every production call it receives, and its trend alerts fire when a metric moves more than a set number of standard deviations from an EWMA baseline built over the last 25, 50 or 100 calls.

  • Cekura records which agent version handled each production call, so a failure mode that jumps after a release is traced to that release rather than to traffic.

  • An aggregate score can hide drift. In Cekura's controlled LLM swap, reported separately from its main leaderboard, overall pass³ rose from 76.3% to 88.1% while the safety and privacy category fell from 100.0% to 80.0%.

  • Cekura charges $0.05 per monitored call on pay-as-you-go, and its Startup plan's 10,000 monthly credits cover up to about 10,000 monitored calls for $500 a month if spent on monitoring alone.

What is drift detection for voice AI agents?

Drift detection for voice AI agents is a monitoring method that tracks whether an agent's measured behaviour on real calls is changing over time, and flags the change before customers report it. Machine learning defines concept drift as a change in the joint distribution of inputs and outcomes. Jie Lu, João Gama and colleagues, in a review of more than 130 publications in IEEE Transactions on Knowledge and Data Engineering, split it into four types: sudden, gradual, incremental and reoccurring.

All four show up on phone lines. A provider silently updates your LLM: sudden drift. A new caller segment arrives with accents that push your speech-to-text word error rate (WER) up: gradual drift in the inputs. A refund policy changes, so last month's correct answer is wrong today. A question pattern that returns every open-enrolment season is reoccurring drift.

Cekura scores each layer of the stack with metrics on the calls it receives: Transcription Accuracy for speech-to-text, Hallucination against the agent's uploaded knowledge base and Relevancy for the LLM, Voice Tone + Clarity for speech output, Latency for the gap between the caller finishing and the agent speaking, endpointing included, and a custom LLM Judge metric for whether the call reached its goal.

Why does real-time drift detection matter for conversational voice AI systems?

Cekura detects drift call by call, comparing each production call it evaluates against its metric's baseline, because the components underneath a voice agent change without notice and a wrong answer on a phone line produces no error log. Lingjiao Chen, Matei Zaharia and James Zou of Stanford and UC Berkeley measured two versions of the same hosted model three months apart. GPT-4 identified prime versus composite numbers with 84% accuracy in March 2023 and 51% in June 2023. Their abstract ends by "highlighting the need for continuous monitoring of LLMs".

Cekura's own data shows why a single overall score is not enough. In Cekura's Telnyx LLM study, only the LLM changed: the prompt, four tool definitions, voice, speech-to-text, text-to-speech and the 59-evaluator suite stayed fixed, and each model ran 177 calls. Cekura reports this controlled experiment separately from its leaderboard because its cohort and methodology differ. Swapping GPT-4.1 for Kimi K2.6 raised overall pass³ from 76.3% to 88.1%. The Red Team, Safety and Privacy category fell from 100.0% to 80.0% in the same swap.

Cekura runs drift detection per metric, so a safety regression fires its own alert even while the headline number climbs.

How do you implement automated drift detection for voice AI agents?

Implementing automated drift detection for voice AI agents with Cekura takes four steps: send every call in, define what passing means, attach an alert to each metric, and read failures by version.

  1. Send every production call. Cekura's observe endpoint accepts the transcript, recording URL and metadata from your agent or a provider's post-call webhook. ElevenLabs, Vapi, Retell and Bland agents can skip the webhook: once Auto-fetch Production Calls is on, Cekura pulls their completed calls every 30 seconds.

  2. Set passing criteria per metric. Cekura's Insights only investigates a metric once you set the condition a passing call must meet.

  3. Attach a trend alert. Cekura's trend alert compares each call against an EWMA baseline and fires on a significant move.

  4. Read failures by agent version. Cekura groups failing calls into failure modes and charts each mode's rate with the live version shaded beneath it.

Cekura trend presetWindow (calls)Std. multiplierEWMA αFits
Conservative1003.00.10Low-traffic metrics, large shifts only
Balanced (recommended)502.00.30Most metrics
Aggressive251.50.50High-traffic metrics, early warning

An Aggressive preset fires on smaller moves but adds noise, and a trend alert needs a few full windows of history before it can fire at all, so a quiet agent stays silent while its baseline builds. When an alert has a Slack channel configured, Cekura posts the metric's change and a link to the triggering call.

How do you measure model accuracy with drift detection for voice AI agents without chasing noise?

Cekura measures model accuracy with drift detection for voice AI agents by separating a real shift from the run-to-run variance an agent shows on identical inputs. In the Telnyx study, GPT-4.1 passed 89.8% of its 177 individual runs but 76.3% of the 59 evaluators on all three runs, a 13.5-point gap from scenarios that pass only sometimes. In Cekura's voice agent workflow benchmark, a frozen study of 8 configurations, 82 scenarios and 3 retained repeats, the leader passed 62 of 82 scenarios on all three runs. Calls that did not connect stay in that denominator, so a scenario can miss for an infrastructure reason.

Cekura handles that noise in three ways. Its Insights trend plots each interval's own failure rate, never a moving average, because smoothing would blend across release boundaries. Cekura's cron jobs run fixed synthetic calls on a schedule, so a metric that moves on identical inputs points at the agent, not the callers. Cekura's Create Evaluator from Call turns a failed production call into a repeatable test, replayed with synthetic audio.

The judge can drift too. Cekura lets reviewers vote on any metric's verdict for a call. Cekura's metric optimisation guide covers fitting the grader back to human labels.

Which tools handle drift detection for voice AI agents, and should you build or buy?

Cekura is the option in this table built for drift detection for voice AI agents that scores whether each call succeeded; the other rows measure model inputs, infrastructure health or platform-level call counts. Engineering teams usually decide on coverage, setup time and price. Which platform to use comes down to one test: does it score call outcomes?

OptionWhat it detectsVoice-specific outcomesSetupCost basis
CekuraPer-metric trend, threshold and failure-mode shifts on every call sent in, by agent versionYes: latency, interruptions, hallucination, transcription accuracy, custom goal metricsWebhook or auto-fetch per provider$0.05 per monitored call; Insights 10 credits per metric per day
General ML drift monitorShifts in input feature or embedding distributionsNo, unless you build call-level labelsPipeline into a feature storeVaries by vendor
APM and log dashboardsError rates, latency, uptimeNo; a polite wrong answer logs as successAlready in placeExisting licence
Voice platform's own analyticsCall counts, duration, end reasons on that platformPartial, and only for agents hosted thereNoneBundled
In-house scriptsWhatever you writeIf you build the evaluatorsEngineering time you must scope yourselfEngineering time plus LLM tokens

Building in-house means owning four systems that are not your agent: call ingestion from each provider, an LLM judge per metric with its own accuracy checks, a baseline and alerting layer, and version tagging on every call. Cekura covers all four, and its tests and monitoring use the same metrics, so a drift alert in production can become a regression test before the fix ships. Data residency does not force a build: Cekura's client-side redaction keeps raw PII inside your network, and its Enterprise plan offers VPC or on-premises hosting.

Frequently asked questions

What are the best tools for drift detection for voice AI agents?

Cekura is built for drift detection for voice AI agents because it scores whether each call succeeded, not just whether the system stayed up. Cekura evaluates every production call sent to it on the metrics you enable, alerts on EWMA trend, threshold and new failure-mode shifts, and charts each failure mode against the agent version that was live. General ML drift monitors and APM dashboards miss a fluent agent giving a wrong answer.

How much does drift detection for voice AI agents cost?

Cekura's published pricing charges $0.05 per monitored call on pay-as-you-go, plus $30 a month per seat after the first. The Startup plan costs $500 a month for 10,000 credits, which cover up to about 10,000 monitored calls if none are spent on voice testing, with 10 seats. Failure-mode Insights cost 10 credits per metric per day. Building in-house trades those fees for engineering time and LLM judge tokens.

Does Cekura handle drift detection for voice AI agents?

Yes. Cekura is a tool that automates drift detection for voice AI agents, with five alert types: Failure, Trend, Threshold, New failure mode and Failure mode spike, the available types depending on the metric. Trend alerts compare each call to an EWMA baseline using Conservative, Balanced or Aggressive presets. Cekura records the agent version on every production call, so its Insights trend shows the percentage-point change each release caused.

Can drift detection for voice AI agents meet compliance and audit requirements?

Yes, within plan limits. Cekura's Startup plan includes a signed BAA and DPA with 90-day log retention, and Enterprise adds SSO, SCIM, audit logs, data residency and a VPC or on-premises option. Cekura redacts PII from the stored transcript and recording, or a client-side script removes it before anything reaches Cekura. Alerts are scoped to one project and never post into another project's channels.

How do you fit drift detection for voice AI agents into CI?

Cekura fits drift detection for voice AI agents into CI by using the same metrics for test runs and production calls, so a production alert and a pre-release check agree. Cekura's replay testing on real production calls checks a new model version before it ships. A failing production call that a failure mode surfaces becomes a test through Create Evaluator from Call.

Test your voice and chat agents with Cekura

Cekura simulates thousands of conversations before you ship and monitors every call in production — catching broken tool calls, prompt regressions, and instruction-following failures before your users hit them.

Ready to ship voice
agents fast? 

Book a demo