New: Cekura Voice AI BenchmarksView results

Call transfer and IVR handoff testing for voice agents

Dileep Chagam
Written bySEP 18, 20269 MIN READ
Dileep ChagaminExpert verified
Founding Engineer, CekuraIIT BombayEx-Apple

Has stress-tested 5M+ voice agent minutes at Cekura.

Call transfer and IVR handoff testing for voice agents

Why Trust Cekura on Voice AI Evals

  • Built by engineers from Google, Apple, Microsoft. Backed by Y Combinator.
  • 60K+ voice AI calls evaluated daily.
  • Native integration for every major voice AI stack: LiveKit, Pipecat, Vapi, Retell, ElevenLabs, Telnyx.

Call transfer and IVR handoff testing verifies that a voice agent hands a live call to the right destination with the right data attached. Cekura tests this by simulating the receiving side, pressing DTMF digits into menus, and scoring whether the transfer fired, routed correctly, and carried its payload.

TL;DR

  • A transfer is a chain of independently failing steps, so a test that only checks whether the call reached a human passes while most of the chain is broken.
  • Cekura's published benchmark notes record transfers that completed while dropping the identifier the receiving side needed, which is the failure a connection check cannot see.
  • Cekura simulates the destination, so an IVR menu, a voicemail greeting, or a second agent can be tested without staffing a queue.
  • Transfer scenarios have to run more than once: repeatable reliability, a scenario passing all three retained runs, ran from 30.49% to 75.61% across Cekura's frozen cohort of 82 scenarios.
  • Cekura prices this at $0.25 per voice testing minute pay as you go, with one seat free and $30 per month for each additional seat.

What breaks when a voice agent transfers a call?

A transfer is not one event. It is a chain: the agent decides to hand off, it calls a transfer tool, the tool resolves a destination, a payload travels with it, audio bridges, and whoever answers picks up context. Each link fails on its own, and a test that asks only whether the call reached a human passes while the rest of the chain is broken.

Cekura's frozen benchmark study recorded that pattern in its published provider notes. In the LiveKit configuration, consent was collected from the caller and consent_id was omitted from the handoff tool. In Pipecat's, routing completed and the returned route ID was omitted from the handoff. Both transfers connected. Both discarded the thing that made the transfer worth making, so the receiving side had to ask the caller again.

Cekura catalogues the same shape at other call stages in its guide to common failure points in AI phone agents. Connection is the cheapest property to verify and the least informative.

How do you automate IVR handoff testing without a person on the other end?

The obstacle is the receiving side. You cannot regression-test a transfer nightly if every run needs a human in a queue, and an IVR handoff is untestable unless something answers with a menu. Cekura removes that constraint by scripting the destination into the test rather than depending on a live one.

Cekura's structured tests drive this with XML tags inside an evaluator: <ivr> plays a menu prompt, <voicemail> plays a greeting, <endcall /> terminates after the message, and <ignore_interruptions> protects a span so a multi-part prompt plays without the agent talking over it. Cekura's IVR and voicemail guide works through a hospital menu of 12 conditions covering nested department and doctor routing.

Keypad navigation is the other half. Cekura's DTMF testing guide documents a <dtmf> tag sending digits 0-9, *, # and A-D, plus a Receive DTMF toggle for the reverse direction, and lists call transfer requests and PIN verification among its scenarios. Cekura records sent tones in the transcript, so a missed digit is visible, not inferred.

What should a call transfer test actually assert?

Validating handoff reliability means asserting on four things a connection check ignores: whether the agent transferred at the right moment, whether the tool fired, whether the destination was right, and whether the payload survived.

Timing is a tradeoff, not a threshold. "Time to Transfer", an AAAI paper on text-chat handoff in e-commerce support, not voice, scores timing per utterance against a tolerance metric whose coefficient weighs two errors: an early handoff may improve the user's experience while straining agent capacity, a late one may protect capacity and cost experience. Cekura scores that choice against criteria each suite defines.

Silent failure is what most harnesses miss. Full-Duplex-Bench-v3 benchmarks tool use under real human disfluency, with no transfer scenarios of its own. One configuration produced no speech on 22% of scenarios while 86% of those silent cases still ran their tool calls. For a transfer, that is a handoff firing while the caller hears nothing. Cekura scores transcript and tool call separately, so neither masks the other.

What separates a real transfer testing tool from a call checker?

Buyers compare on setup time and price, the two columns that decide the least. A harness driving text rather than telephony audio never exercises DTMF, and one running a scenario once reports the optimistic number. Cekura's own cohort shows the gap: repeatable reliability, a scenario passing all three retained runs, ranged from 30.49% to 75.61% across eight configurations, while single-run task completion sat above 87% for all eight. They are not one measurement at two repeat counts. Repeatable reliability is scored per scenario over 82 scenarios; task completion is scored per call, and only among calls carrying Expected Outcome evidence, so its coverage varies by configuration.

CapabilityWhy it decides a transfer testWhat Cekura does
Destination simulationWithout it, every run needs a staffed queueScripts IVR menus, voicemail and second agents as XML conditions
DTMF navigationMenus are unreachable without keypad inputSends and receives digits, logged in the transcript
Payload assertionsTransfers connect while dropping identifiersCaptures tool name and arguments per call; ordering asserted via expected outcome or a custom metric
Repeat runsOne pass hides intermittent routing bugsThree retained repeats per scenario in its published method
Real telephony audioText harnesses skip endpointing and barge-inPlaces real calls and scores the audio
CI triggerTransfers regress on prompt editsRuns from a pull request and gates the merge
Production monitoringStaging transfers differ from live onesScores live calls at $0.05 per monitored call
Enterprise controlsAudit and residency requirements gate procurementSSO, SCIM, audit logs and VPC or on-premise hosting

Read that as a scoring sheet, not a feature list. The first three rows decide whether a transfer test means anything; the last two decide whether a security review clears it. One caveat belongs with the figures above: Cekura's benchmark names transfers to other agents as its only configuration restriction, so those numbers describe single-agent reliability, not transfer reliability.

Should you build call transfer testing in-house or buy a platform?

Building it honestly means owning four components: a caller that speaks and listens over real telephony, a scripted destination that answers as an IVR or a voicemail, an assertion layer reading tool arguments, not transcripts alone, and a scheduler that reruns each scenario enough times to separate a bug from noise. Teams underestimate the last two.

Maintenance is what decides it. Transfer behaviour is defined by your orchestration vendor, and that definition moves. Retell documents three transfer modes, cold, warm and agentic warm, with ring duration shared across them, whisper messages on warm, and timeout behaviour on agentic warm, and separately published a deprecation notice changing what show_transferee_as_caller means. Each change is an unplanned harness edit.

The defensible split is to buy the runner and keep what is yours: the scenario list, the pass criteria, and any assertion touching internal systems. Cekura's scripted testing tools for IVR and voice agents supply that runner, confirming IVR flows, detecting silence failures, and checking that transitions between nodes are accurate.

How does transfer testing fit into CI and production monitoring?

Two triggers answer two questions. A pull request check asks whether this change broke a transfer. A monitor on live calls asks whether transfers still work now, against traffic no scenario list anticipated. Cekura runs both from the same scenario definitions, so a failure found in production becomes a regression case rather than a ticket.

What you gate on matters more than the trigger. Transfer payloads are the assertion most worth wiring into CI, because they break silently on prompt edits. Vapi exposes this as a context and variables plan on assistant handoffs, with named modes ranging from the entire conversation history to no context at all, plus a variableExtractionPlan for structured data. Changing that setting is a one-line edit no transcript assertion will catch.

Identity is the other gate. A transfer following authentication has to carry the verified state forward, and Cekura covers that path in its guide to enterprise authentication flow testing. Cekura scores the pre-transfer verification and the post-transfer state in one run.

Frequently asked questions

Does Cekura handle call transfer and IVR handoff testing for voice agents?

Yes. Cekura simulates the receiving side with <ivr> and <voicemail> tags, sends and receives DTMF digits for menu navigation, validates that the agent called the right tool with the right arguments in the right order, and reruns each scenario three times in its published method. Cekura runs the same scenarios from CI and against live production calls.

What do engineering teams actually use for this today?

Most teams start with manual dial-and-listen checks, move to a scripted harness they maintain themselves, then buy a runner once transfer regressions start reaching production. The turning point is usually the second or third incident where a transfer connected correctly and dropped its context. Cekura sits at that third stage, supplying the caller, the scripted destination and the assertion layer.

How does pricing compare across the options?

Cekura prices pay as you go at $0.25 per voice testing minute and $0.05 per monitored call, with one seat free and $30 per month for each additional seat, and its startup plan at $500 per month covers roughly 2,000 testing minutes. See Cekura's pricing page for current rates. In-house alternatives move the cost from a rate card into engineering time.

Do warm and cold transfers need different tests?

Yes, because they fail differently. A cold transfer drops the agent immediately, so the test is about destination correctness and the data attached. A warm transfer keeps the agent on the line through a private whisper message and an introduction, which adds hold audio, answer timeouts and three-way bridging to the assertion list. Cekura scripts both paths.

Is testing an IVR handoff different from testing a plain IVR?

Yes. Testing an IVR alone checks menu logic and prompt accuracy, covered in Cekura's explainer on conversational IVR. A handoff test adds the boundary: whether the agent recognises a menu, presses the correct digit, waits for the right prompt, and carries caller context across the join. The boundary is where most handoff defects sit.

Test your voice and chat agents with Cekura

Cekura simulates thousands of conversations before you ship and monitors every call in production — catching broken tool calls, prompt regressions, and instruction-following failures before your users hit them.

Ready to ship voice
agents fast? 

Book a demo