Using Claude Code, Cursor, or Codex to run voice agent tests means connecting the coding assistant to a test platform over MCP, so it can create scenarios, start simulated calls and read scored transcripts without leaving your repo. Cekura ships one plugin for all three that bundles its MCP server with Cekura Skills.
TL;DR
- The coding assistant does not place calls itself. Cekura runs the simulated caller, dials or joins the agent, records the call and scores the transcript, while the assistant drives it through MCP tools.
- Cekura's plugin installs in Claude Code, Cursor and Codex, signs in over OAuth with no API key, and gives the assistant Cekura Skills covering evaluator design, metric design, CI suites and failure triage.
- Cekura's Tests as Code endpoint takes a committed
cekura.tests.jsonfile and, withdry_run=true, returns the case count and estimated cost without placing a single call. - Run each scenario more than once. In Cekura's frozen benchmark the leading configuration passed 62 of 82 scenarios on all three retained runs, counting calls that did not connect as failures.
- Cekura charges $0.25 per voice testing minute and $0.025 per text reply, and lists the API, MCP server and Skills as included on every plan.
What does it mean to use Claude Code, Cursor, or Codex to run voice agent tests?
Using Claude Code, Cursor, or Codex to run voice agent tests is a workflow in which the coding assistant acts as the operator of a voice test platform. The assistant reads your prompts and tool definitions, decides what to test, and calls the platform's tools. The platform does the part a coding assistant cannot: it dials or joins the agent, plays a simulated caller, records the audio and scores what was said. This is not speaking to the assistant by voice; the voice belongs to the conversational AI agent under test.
Per the Cekura MCP and Skills setup guide, Cekura's MCP server lets an assistant list agents, create metrics, generate evaluators, run tests, inspect transcripts and analyze results. Cekura splits the run step by connection: each scenarios_run_* tool drives one transport: scenarios_run_voice for telephony, scenarios_run_sip, scenarios_run_vapi_webrtc, scenarios_run_retell_webrtc and scenarios_run_text for chat, with further tools for ElevenLabs, LiveKit, Pipecat, Agora and raw WebSocket agents.
Cekura Skills add judgment: the MCP server supplies tools, and the skills tell the assistant how to design an evaluator, pick a metric and read a failed run. For the raw endpoints underneath, see Cekura's programmatic voice agent testing API.
How do you set up Claude Code, Cursor, or Codex to run automated voice agent tests?
Setting up Claude Code, Cursor, or Codex to run automated voice agent tests with Cekura is one plugin install and one OAuth sign-in per assistant. All three pull from the same cekura-ai/cekura-skills marketplace, so the skills and the MCP configuration match across a team that uses different editors.
| Assistant | Install | Sign-in | How you trigger a test run | Updates |
|---|---|---|---|---|
| Claude Code | /plugin marketplace add cekura-ai/cekura-skills, then /plugin install cekura@cekura-skills | /setup-mcp, OAuth in the browser | Slash commands such as /run-evals and /eval-results, or plain English | /upgrade-skills, or marketplace auto-update |
| Cursor | Settings, Plugins, Team Marketplaces, Add Marketplace, Import from Repo | OAuth when prompted | Plain English; skills activate when the request matches | Re-import, or Auto Refresh through the Cursor GitHub App |
| Codex | codex plugin marketplace add cekura-ai/cekura-skills, then codex plugin add cekura@cekura | codex mcp login cekura | Plain English, or invoke a skill with @ | A bundled SessionStart hook, trusted once |
| Any other Agent Skills client | npx skills add cekura-ai/cekura-skills --all | Separate MCP setup | Plain English | npx skills update, run by hand |
Cekura runs its MCP server in three regions, api.cekura.ai, api.in.cekura.ai and api.eu.cekura.ai, so a workspace hosted in India or the EU points the assistant at its own endpoint. Cekura recommends OAuth for people and an API key only for CI or project-scoped credentials. To confirm the connection, ask the assistant to list your Cekura agents and propose three evaluators without creating anything. A working setup returns real agents from your workspace.
Why should the coding agent reproduce a voice agent failure as a test before fixing it?
Cekura's cekura-self-improving-agent skill makes the coding agent reproduce a failure in Cekura simulation before it proposes any edit. A reproduction test fails on the current code and passes once the fix lands, which proves the patch fixed the reported problem.
The evidence comes from code repair. Niels Mündler, Mark Niklas Müller, Jingxuan He and Martin Vechev of ETH Zurich and LogicStar.ai built SWT-Bench, a benchmark of real GitHub issues, and had code agents write tests that reproduce each issue. Keeping only the proposed fixes whose generated tests all went from failing to passing or stayed passing "more than doubles the precision of SWE-Agent to 47.8%", at a recall of 20%. The study covers Python repositories, not voice agents, and the tradeoff is recall: the filter discards most candidate fixes.
Cekura's documented Codex workflow applies the same order to voice. Run the assistant from the agent's repo, approve its coverage plan, run a small batch, then have it compare the failed transcript against the prompt or code, patch the likely cause and rerun the one evaluator that failed. For suites that must keep passing after every change, see regression testing for voice AI agents.
How does a coding agent keep voice agent tests in the repo and fit them into CI?
A coding agent keeps voice agent tests in the repo by writing them as a JSON spec that CI submits to Cekura, not as dashboard records. Cekura calls this Tests as Code: a cekura.tests.json file lists each scenario's caller instructions, expected outcome, metrics, personality and test data, and nothing in it becomes an evaluator in your workspace.
Cekura prices a spec before spending anything: posting it with dry_run=true returns the scenario count, total runs and estimated cost, and creates nothing. Dropping the flag starts the runs, on a voice, text, elevenlabs, livekit_v2 or pipecat_v2 channel, and results land in Results like any other run.
Two of the 13 skills in the cekura-skills repository do the repo work. cekura-infra-test-suite inspects the code, writes the spec, a .github/workflows/cekura-tests.yml gate and a README section, and treats Cekura as read-only apart from validation. cekura-bot-test-writer reads a pull request's diff against its merge base and makes the smallest edit that covers a real gap. Its stated default is no change, because every added case is a live call. For what the merge gate itself should check, see CI/CD testing for voice AI agents.
Which tools let Claude Code, Cursor, or Codex run voice agent tests, and should you build or buy?
The tools that let Claude Code, Cursor, or Codex run voice agent tests differ on the question engineering teams weigh before setup time and price: whether a test reaches a real call through speech recognition, turn-taking and synthesis, or only exercises the agent's logic in text.
| Option | What the assistant can do | Reaches a voice call | Where tests live | Cost basis |
|---|---|---|---|---|
| Cekura plugin (Skills plus MCP) | Design, run, triage and patch, with workflow guidance | Yes, plus a text channel | Dashboard evaluators or a committed JSON spec | $0.25 per voice testing minute, $0.025 per text reply |
| Cekura MCP server only | Call the same tools, without the skills' guidance | Yes | Dashboard or JSON spec | Same usage rates |
| Cekura Skills only | Advise and write specs, with no live tool access | No, until MCP is connected | Your repo | No usage until a run |
| Unit tests the assistant writes itself | Check prompt logic and tool arguments in text | No | Your repo | Your LLM tokens |
| An in-house MCP server over your own harness | Whatever you build | If you build a caller | Your repo | Engineering time plus telephony |
Building the in-house row means owning five systems that are not your agent: a caller that dials or joins a room, an LLM that plays the customer, graders per metric, a runner that survives long calls, and storage for audio a reviewer can open. Each change to your telephony provider or speech stack has to be mirrored in that harness. It wins when your stack has no connector or your test data cannot leave your network. Cekura replaces the harness with hosted callers and evaluators behind the same MCP tools. As an illustration, 12 scenarios run 3 times at an assumed 2 minutes a call is 12 × 3 × 2 = 72 voice minutes, or $18 at Cekura's published rate.
What are the best practices for using Claude Code, Cursor, or Codex to run voice agent tests?
The best practices for using Claude Code, Cursor, or Codex to run voice agent tests control two things a coding agent will otherwise waste: credits and context.
- Start read-only. Have the assistant list agents, results and call logs before it creates or runs anything.
- Approve the plan before anything is created. Ask for a coverage matrix of workflow, edge-case, regression and tool-call cases first.
- Dry-run before spending. Validate a spec with
dry_run=trueand read the estimated cost before any live call. - Scope and paginate. Pass
project_idacross several projects, and page large result sets so transcripts do not fill the context window. - Run each scenario more than once. Cekura's voice agent workflow benchmark is a frozen study of 8 configurations, 82 scenarios and 3 repeats. Its leader passed 62 of 82 scenarios on all three retained runs, so 20 failed at least one. Calls that did not connect count as failures there, so the 20 mix repeatable, intermittent and infrastructure failures. Start at a frequency of one, then raise it once the first run succeeds.
Cekura records every run in Results with its transcript and metric scores, so a reviewer can open exactly what the assistant tested.
Frequently asked questions
What are the best tools for using Claude Code, Cursor, or Codex to run voice agent tests?
Cekura is the tool that places real simulated calls and exposes them to the assistant over MCP; unit tests the assistant writes itself only check logic in text. Cekura ships a single plugin for Claude Code, Cursor and Codex that bundles its MCP server with Cekura Skills, and it runs scenarios over telephony, SIP, WebRTC, LiveKit, Pipecat and a text channel.
How much does it cost to run voice agent tests from Claude Code, Cursor, or Codex?
The cost is the test platform's usage, since the assistant only orchestrates. Cekura charges $0.25 per voice testing minute and $0.025 per text reply, starts accounts with 300 free credits (about 60 minutes), and on pay-as-you-go includes one seat free with $30 a month per additional seat. The Startup plan includes 10 seats, and the pricing page lists the MCP server and Skills on every plan. Multiply scenarios by repeats by average call length to estimate a run.
Does Cekura handle using Claude Code, Cursor, or Codex to run voice agent tests?
Yes. Cekura publishes a plugin for Claude Code, Cursor and Codex from the cekura-ai/cekura-skills repository, which is MIT-licensed. It installs Cekura Skills and connects the Cekura MCP server over OAuth. Claude Code also gets slash commands such as /run-evals, /eval-results and /autogen-eval. GitHub Copilot and Gemini CLI have install paths too.
Can a coding agent run voice agent tests under compliance and audit requirements?
Yes, within the plan's limits. Cekura's Startup plan lists a signed BAA and DPA and 90-day log retention, and the Enterprise plan adds SSO, SCIM, audit logs, access control across projects and a VPC or on-prem option. Keep API keys in CI secrets, never in a committed spec, and use OAuth for people at the keyboard.
Does Claude Code place real phone calls when it runs a voice agent test?
No. Claude Code, Cursor or Codex calls a Cekura MCP tool, and Cekura places the call. The tool the assistant picks decides the transport, so scenarios_run_voice dials over telephony while scenarios_run_text runs a chat conversation. Cekura bills voice runs per testing minute, which is why a dry run that returns the estimated cost comes first.







