New: Voice AI Orchestration Benchmarks — Retell, Vapi, Pipecat, LiveKit & more

SOC2 compliant voice AI testing platform

Satvik Dixit
Written bySEP 4, 20269 MIN READ
Satvik DixitinExpert verified
Founding Engineer, CekuraMS, CMU

Has stress-tested 5M+ voice agent minutes at Cekura.

SOC2 compliant voice AI testing platform

Why Trust Cekura on Voice AI Evals

  • Built by engineers from Google, Apple, Microsoft. Backed by Y Combinator.
  • 60K+ voice AI calls evaluated daily.
  • Native integration for every major voice AI stack: LiveKit, Pipecat, Vapi, Retell, ElevenLabs, Telnyx.

A SOC2 compliant voice AI testing platform holds an independent SOC 2 Type II report covering the systems that ingest your call recordings, transcripts and prompts. Cekura holds SOC 2 Type II alongside HIPAA and GDPR compliance, and runs inside your own cloud when those recordings cannot leave it.

TL;DR

  • A SOC 2 Type II report describes the vendor's controls over a stated observation window. Cekura holds one. It is still not a certificate about your agent, and it says nothing about whether your calls are lawful.
  • The report only matters where its scope reaches. Ask whether the audited scope includes the call ingestion path, the recording store and the evaluation pipeline, not just the marketing website.
  • A testing platform is a higher risk than most vendors because it holds raw audio. Speaker identity survives content redaction, which is why Cekura treats storage location and access control as the primary control and redaction as a second layer.
  • PCI work is a separate question with a separate answer. Cekura redacts card data before storage, and the telephony layer has its own PCI switches that can disable the transcription your tests read.
  • Cekura offers deployment inside your own cloud and regional data residency across the US, Europe and India, so the audited boundary can sit where your policy requires.

What does it mean for a voice AI testing platform to be SOC 2 compliant?

A SOC 2 compliant voice AI testing platform is one whose controls have been examined by an independent auditor against the AICPA Trust Services Criteria, and which can hand you that report. Cekura holds a report of that kind, a SOC 2 Type II. Two things in that sentence do the work, and buyers skip both.

The first is the report type. A Type 1 report says controls were suitably designed at one point in time. A Type II report says they also operated effectively across an observation window, and it names its own start and end dates. A Type 1 in a procurement pack answers a weaker question than you asked.

The second is scope. SOC 2 has no fixed boundary: the vendor defines which systems and criteria the examination covers. Cekura holds SOC 2 Type II certification, HIPAA compliance and GDPR certification, as its self-hosted voice agent testing page states. Read any vendor's report for what sits inside that boundary before treating the badge as an answer.

Why does the testing platform's own security posture matter more than the agent's?

A voice AI testing platform is unusual among vendors because it holds raw call audio, full transcripts and your system prompts in one place. That is a richer set than most subprocessors see, so the vendor's posture deserves more scrutiny than a badge check.

Content redaction does not solve it, because a speaker's identity lives in the voice itself and not only in the words. A 2025 benchmark reports that "all state-of-the-art speaker de-identification systems leak identity information", with the weakest system reaching a 45% hit rate within the top 50 candidates. Read that as research on voice anonymisation, not on any redaction product: it measures systems built to disguise speakers, not transcript scrubbers. It transfers anyway: a recording stays identifying, so where it lives and who reaches it is the control that matters.

Cekura therefore treats the storage boundary as the primary control and redaction as a second layer. Which rules apply to those calls is separate, covered in our AI voice agent compliance guide.

What should you check in a vendor's SOC 2 report before you buy?

Ask for the report itself, under NDA, and read four things: the report type, the observation window, the systems inside scope, and the exceptions the auditor recorded. Exceptions are the most useful page in the document and the one most often skipped, because a report with none is either a narrow scope or a short window.

CriterionWhy it decides the purchaseCekura
SOC 2 Type II report, released under NDAA Type 1 or a summary page answers a weaker questionYes, SOC 2 Type II
Audited scope reaches the call ingestion pathThe badge is worthless where the recordings actually sitConfirm against the report
Deployment inside your own cloudSome recordings cannot leave your boundary at allYes, runs in the customer cloud
Regional data residencyResidency is a policy constraint, not a preferenceUS, Europe and India
Redaction before storageReduces what a breach of the vendor would exposeYes, entity-level, before storage
Named read-only role for auditorsAuditors need evidence, not administrative accessViewer role and read-only API keys

Demand the same discipline of any evidence a vendor publishes about itself. Cekura's voice agent benchmarks run 7 configurations over 82 scenarios with 3 retained repeats, and state that "Calls that did not connect or produced no transcript stay in the denominator" for infrastructure reliability, while task completion counts only calls carrying outcome evidence. Those are platform performance figures rather than security results, and they are cited here for the method: a vendor that shows where each metric draws its denominator is showing you how it counts.

How is PCI compliant voice agent testing different from SOC 2?

PCI compliant voice agent testing asks a narrower question than SOC 2: whether cardholder data is kept out of everything your test pipeline stores. Cekura answers it by removing sensitive entities before storage. SOC 2 covers the vendor's control environment; PCI DSS constrains one data type wherever it goes.

Two layers cooperate. On the testing side, Cekura removes sensitive entities from the stored transcript and recording, working from an explicit list that includes card numbers, social security numbers and healthcare numbers, per its PII redaction documentation.

The telephony side is less obvious and breaks tests. Twilio's own guidance states that "Native and Marketplace transcriptions are not available when PCI Mode is enabled", directing customers to the Voice Transcription noun instead. That path survives PCI Mode, so a PCI-configured account can quietly cut the one your evaluation layer was wired to. Find that in procurement, not in a release. Whether the agent behaves correctly on a payment turn is separate, covered in our compliance testing for voice AI agents guide.

Where should the call data live, and who should hold the detection keys?

Data location is the decision that survives every audit cycle, so settle it before the security questionnaire. Cekura runs inside a customer's own cloud, and offers regional data residency across the US, Europe and India, where databases, storage, caches and processing stay inside the selected geography.

For teams whose raw audio must never leave the network, Cekura supports detection on your own infrastructure. Redaction runs in your own process, on your own language and speech providers, with your own API keys, and the documented failure behaviour is the part worth noting: the client-side redaction path fails closed, raising an error rather than returning unredacted content when detection or alignment breaks down. A pipeline that fails open quietly ships raw card numbers into a transcript store.

The surrounding practices are ordinary and still get skipped: encryption in transit and at rest, short retention windows with automatic deletion, and role scoping that keeps read access narrow. Our data privacy best practices guide works through those controls for voice.

Should you buy a SOC 2 compliant testing platform or build one in-house?

Build if your call volume is low, your data cannot leave one region, and you already run an audited platform to extend. In that position the incremental control work is small and you inherit an existing scope. Buy when the same suite must run across several agents and platforms, or when somebody outside engineering needs an evidence trail.

The honest cost of building is not the harness. It is the security surface you take on: an audio store that becomes in-scope for your own audit, a redaction pipeline you maintain and calibrate as models change, key management for the detection providers, and a retention policy somebody enforces. A testing pipeline that accumulates production recordings is a data liability with a test-tool budget.

Price the second year, not the first. Most teams that build a voice testing harness choose to expand their own audit scope, which is defensible as long as it was a choice. Book a demo to see the deployment and redaction paths against your own calls.

Frequently asked questions

What are the best tools for a SOC2 compliant voice AI testing platform?

Judge candidates on four checks: whether they will release a SOC 2 Type II report under NDA, whether the audited scope covers the recording store, whether they can deploy inside your cloud, and whether redaction runs before storage. Cekura meets all four. Tools built for offline model evaluation frequently fail the third, because they were never designed to hold live call audio.

Which platform should I use for SOC2 compliant voice AI testing?

Start from your hardest constraint rather than the feature list. If audio cannot leave your network, the deployment model decides it and most hosted evaluation tools are eliminated immediately. If residency is the constraint, the region list decides it. Cekura supports both the customer-cloud deployment and regional residency, which is why regulated teams tend to shortlist on those two answers first.

How do platform options compare on pricing for compliance-heavy testing?

Compliance suites are small in scenario count and high in run frequency, so usage pricing usually suits them better than per-seat pricing. Cekura meters voice testing by the minute and monitoring by the metric run, and charges seats separately from that usage, so a suite that runs often costs more than one that runs wide. Ask any vendor how repeat runs and a customer-cloud deployment are priced before you compare totals, since either can move the number more than the per-minute rate suggests.

Can a SOC2 compliant voice AI testing platform run inside our CI pipeline?

Yes, and it should. Cekura triggers suites from CI so a prompt change is scored before it reaches a campaign, with per-scenario results retained as the evidence trail. The security question CI raises is credential scope: use a project-scoped or read-only API key in the pipeline rather than an administrative one, and rotate it on the same schedule as your other build secrets.

Does Cekura handle SOC2 compliant voice AI testing?

Yes. Cekura holds SOC 2 Type II certification, HIPAA compliance and GDPR certification, deploys inside a customer's own cloud, offers regional data residency across the US, Europe and India, redacts sensitive entities before storage, and supports detection on your own infrastructure with a fail-closed path. Cekura does not certify your agent or your programme as compliant.

Test your voice and chat agents with Cekura

Cekura simulates thousands of conversations before you ship and monitors every call in production — catching broken tool calls, prompt regressions, and instruction-following failures before your users hit them.

Ready to ship voice
agents fast? 

Book a demo