Self-hosted voice agent testing is automated testing of a voice agent running on infrastructure you operate, such as a self-hosted LiveKit or Pipecat deployment. Cekura tests these agents over SIP, WebRTC or a raw-audio WebSocket, reaches private SIP endpoints from one allowlisted egress IP, and lists VPC and on-prem deployment on its Enterprise plan.
TL;DR
- Self-hosted voice agent testing is two decisions, not one: where the agent runs, and where the test harness and its recordings run.
- Text-mode tests overstate voice performance. In the tau-Voice benchmark, a reasoning text model completed 85% of 278 tasks while separate audio-native voice agents completed 31-51% on clean audio and 26-38% with noise and accents.
- One run per scenario hides flakiness. In Cekura's frozen benchmark, the LiveKit and Pipecat configurations passed 70.73% and 63.41% of 82 scenarios on all three runs, against 95.12% and 94.21% task completion counted only among calls with Expected Outcome evidence.
- Cekura reaches an agent on a private network over SIP from a single static egress IP, and its client-side redaction strips PII inside your network before anything is sent.
- Cekura's Pay as you go tier costs $0.25 per voice testing minute, with one free seat and $30 a month for each additional seat, as published in September 2026.
What does self-hosted voice agent testing actually cover?
Self-hosted voice agent testing is a practice covering two separate choices: where the agent under test runs, and where the test harness, recordings and transcripts live. Teams that self-host the agent, on LiveKit or Pipecat in their own cloud, then have to decide whether the harness must follow it.
The two choices rarely need the same answer. The real requirement is usually a data boundary: which audio and personal data may leave your network. Cekura handles the common case, a self-hosted agent tested by a hosted harness, through private-network connectivity and redaction. For teams whose boundary rules out a hosted harness, Cekura lists VPC and on-prem deployment on its Enterprise plan.
Framework-native tools do not settle the question. LiveKit's testing documentation states that its agent simulations "run on LiveKit Cloud through the authenticated CLI", so a self-hosted LiveKit server still sends simulated conversations to a hosted service. For network and hardware requirements, see Cekura's guide to on-premise voice AI testing deployment.
Which options exist for self-hosted voice agent testing, and how do they compare?
Self-hosted voice agent testing options fall into four groups: a harness you build, the test tools shipped with your agent framework, a hosted testing platform, and a platform deployed inside your own cloud. They differ most on whether they exercise real audio and on who maintains the harness.
| Criterion | Build in-house | Framework test tools | Hosted platform | Cekura |
|---|---|---|---|---|
| Where the harness runs | Your servers | Your CI, plus vendor cloud for simulations | Vendor cloud | Cekura cloud, or VPC and on-prem on Enterprise |
| Real audio path tested | Only if you build a media client | Text by default, audio optional | Varies by vendor | SIP, WebRTC and raw-PCM WebSocket |
| Reaches an agent behind a firewall | Yes | Local text tests only | Varies | SIP from one allowlisted egress IP |
| Repeats per scenario | You build it | You build it | Varies | Configurable frequency per evaluator |
| CI trigger | You build it | Native | Varies | GitHub Action and a tests-as-code JSON spec |
| Who maintains it | You | Framework maintainers | Vendor | Cekura |
Choose on the audio row first. A harness that tests only text leaves endpointing, barge-in and speech recognition errors untested, and those are the failures a self-hosted media stack adds. Cekura's comparison of how to choose a voice AI testing platform covers the wider criteria.
Cost follows the maintenance row. A built harness carries no licence fee but a standing engineering cost in media handling, caller simulation and scoring. Framework tools carry no separate licence and stop at text by default. Hosted platforms and Cekura bill for usage, which moves that maintenance to the vendor.
Why do text-mode tests miss what breaks a self-hosted voice agent?
Text-mode testing is a method that sends the agent typed turns instead of audio, so it tests prompt, tools and policy while skipping speech recognition, turn-taking and synthesis. LiveKit's documentation calls text mode "the most cost-effective and deterministic way to test agent behavior" and advises teams to "reserve audio runs for turn-taking and speech-specific issues".
Those speech issues are large. In tau-Voice, a March 2026 benchmark from Sierra and Princeton University, a reasoning text model completed 85% of 278 grounded customer-service tasks. Audio-native voice agents completed 31-51% on clean audio and 26-38% with noise and diverse accents. These are different models, so read the gap as the size of the voice penalty, not a like-for-like score.
Self-hosting widens what audio tests must cover: codec negotiation, packet handling and endpointing settings become yours to break. Cekura runs each scenario as a real call over SIP, WebRTC or WebSocket, and scores latency, interruptions and infrastructure issues alongside the transcript.
How does a test harness reach a voice agent behind your firewall?
A test harness reaches a self-hosted voice agent through one of three transports: SIP, WebRTC, or a direct audio WebSocket. Each needs an inbound path your network allows.
For SIP, Cekura sends a SIP INVITE carrying X-Run-Id, X-Scenario-Id and X-Result-Id headers, so your logs map each call to its test run. Per Cekura's IP whitelisting documentation, all production traffic egresses through one NAT gateway at 32.184.152.201 in US-WEST-2, and that address is the one to allow for "SIP endpoints hosted on private networks or behind firewalls".
For WebRTC, Cekura takes your LiveKit server URL, API key and secret, then creates a room and token per scenario and cleans up. A self-hosted LiveKit server must expose signalling and media ports: LiveKit's deployment guide configures port 7880, TCP 7881 and UDP 50000 to 60000.
For custom stacks, Cekura dials a wss:// endpoint directly and exchanges 16 kHz mono PCM frames, with optional Basic Auth on the upgrade request.
How do you scale self-hosted voice agent testing across many scenarios and keep monitoring it?
Scaling self-hosted voice agent testing means running many scenarios, several times each, on every change, then on schedule. Cekura's load testing guide does this through a frequency parameter: ten evaluators at frequency five place 50 concurrent calls, within your plan's concurrency limit.
Repeats matter because a scenario that passes once can fail intermittently. Per Cekura's frozen benchmark of 8 configurations, 82 scenarios and 3 retained repeats, the LiveKit configuration passed 70.73% of scenarios on all three runs and Pipecat 63.41%. Their task completion was 95.12% and 94.21%, counted only among calls with Expected Outcome evidence. Providers chose their models and settings, so these describe provider-submitted configurations, not your own self-hosted deployment.
For CI, Cekura publishes a GitHub Action that selects scenarios by ID or tag, and a tests-as-code JSON format whose dry run returns estimated cost. After release, Cekura's cron jobs place scheduled calls and alert on failures, and a custom stack can post transcripts to a per-project webhook.
How do you keep test and monitoring data inside a secure environment?
Keeping test data inside a secure environment means controlling three things: what leaves your network, who in the vendor workspace can read it, and where it is stored. Cekura has a separate control for each.
For what leaves, Cekura's client-side redaction detects and removes PII in your own process, on your own LLM and speech-to-text providers, so "the sensitive values never reach Cekura at all". It fails closed, raising an error rather than returning content when a span cannot be aligned to the audio. Server-side redaction costs 0.4 credits per minute.
For access, Cekura's Viewer role cannot open raw transcripts or recordings, and Members see only their assigned projects, so dev, staging and production can sit in separate projects.
For storage, Cekura documents EU residency for production monitoring data, while synthetic testing data is currently hosted in the US. Teams that need everything in their own account can use Cekura's Bring Your Own Cloud deployment, which runs entirely within the customer's own cloud environment.
Frequently asked questions
Should you build self-hosted voice agent testing in-house or buy a platform?
Build in-house only if you will own a media client, a caller simulator, an evaluation layer and a repeat scheduler, and keep them working as your agent changes. That is a standing engineering commitment, not a one-off project. Cekura supplies all four, reaches self-hosted agents over SIP, WebRTC or WebSocket, and moves the maintenance to Cekura, at the cost of a usage bill.
How much does self-hosted voice agent testing cost with Cekura?
Per Cekura's pricing page, as published in September 2026, Pay as you go costs $0.25 per voice testing minute and $0.05 per monitored call, with 10 concurrent calls, one project, one free seat and $30 a month for each additional seat. The Startup Plan is $500 a month with 50 concurrent calls and 5 projects. VPC and on-prem deployment are on the custom-priced Enterprise plan.
Which tools do engineering teams use for self-hosted voice agent testing?
Engineering teams typically combine their framework's text-mode tests for fast iteration with an audio-path harness for turn-taking and speech failures. Cekura fills the second role: it runs scenarios as real calls against self-hosted LiveKit, Pipecat, SIP or WebSocket agents, repeats each scenario at a configurable frequency, and triggers from CI through a GitHub Action.
Does Cekura handle self-hosted voice agent testing?
Yes. Cekura tests self-hosted agents over SIP with test metadata in custom headers, over WebRTC using your own LiveKit server URL, and over a raw 16 kHz PCM WebSocket. It reaches private SIP endpoints from one allowlisted egress IP, supports client-side PII redaction inside your network, and lists VPC and on-prem deployment on its Enterprise plan.







