New: Cekura Voice AI BenchmarksView results

AssemblyAI Pricing: Plans, Costs & What You Get (2026)

Rishabh Sanjay
Written bySEP 25, 202615 MIN READ
Rishabh SanjayinExpert verified
Founding AI Engineer, CekuraMS CS, PurdueEx-Oracle

Has stress-tested 5M+ voice agent minutes at Cekura.

Why Trust Cekura on Voice AI Evals

  • Built by engineers from Google, Apple, Microsoft. Backed by Y Combinator.
  • 60K+ voice AI calls evaluated daily.
  • Native integration for every major voice AI stack: LiveKit, Pipecat, Vapi, Retell, ElevenLabs, Telnyx.

AssemblyAI pricing starts at $0.15 per hour for Universal-2 transcription, but that rate covers one of eight priced products.

Streaming, the Sync API, the Dictation API, and the Voice Agent API each carry their own rates, and add-ons like diarization, medical mode, and PII redaction bill separately on top.

We went through AssemblyAI's full rate card, product by product, to see what a real workload actually costs once the add-ons stack up.

This guide breaks down what each product costs today, which add-ons change the bill, and how AssemblyAI pricing compares to that of Deepgram and Gladia.

TL;DR: AssemblyAI pricing

  • Pre-recorded transcription starts at $0.15/hr for Universal-2. The flagship Universal-3.5 Pro model runs $0.21/hr.
  • Real-time streaming starts at $0.15/hr, but the newer Universal-3.5 Pro Realtime model costs $0.45/hr.
  • The Voice Agent API is a flat $4.50/hr ($0.075/min) covering STT, LLM, TTS, and hosting, billed on open session time. LLM Gateway calls made by the agent bill separately.
  • Add-ons like diarization, medical mode, and PII redaction bill separately, per hour, on top of the base model rate.
  • The free tier is $50 in credits, worth up to 185 hours of pre-recorded or 333 hours of streaming transcription at base rates, with no card required.

AssemblyAI Pricing at a Glance

πŸ›  ProductπŸ’° Price🎯 Best For
Universal-2 (pre-recorded)$0.15/hrBudget batch transcription
Universal-3.5 Pro (pre-recorded)$0.21/hrHighest-accuracy async work across 18 languages
Universal-Streaming (real-time)$0.15/hrCost-sensitive, English-only live transcription
Universal-Streaming Multilingual (real-time)$0.15/hrLow-cost live transcription in English, Spanish, German, French, Portuguese, and Italian
Universal-3.5 Pro Realtime$0.45/hrProduction voice agents needing top accuracy
Sync API$0.45/hrSingle-call, low-latency transcription (~134ms p50)
Dictation API$0.62/hrShort spoken input returned as ready-to-send text
Voice Agent API$4.50/hr ($0.075/min)Fully managed voice agent, all layers bundled
Custom/EnterpriseContact salesHigh-volume, custom concurrency or SLA needs

Pricing verified against assemblyai.com/pricing on September 8, 2026. Verify with AssemblyAI before you budget.

AssemblyAI Pricing Plans Breakdown

AssemblyAI doesn't sell subscription tiers. Every product bills by the hour (or the minute, for the Voice Agent API), and most of them carry their own add-on menu. Here's what each one actually includes.

Universal-2: $0.15/hr

What's included: AssemblyAI's value-tier async model, trained on over 12.5 million hours of audio and supporting 99 languages. Speaker diarization adds $0.02/hr, and medical mode adds $0.15/hr. Free-form prompting isn't supported on this model. Keyterms prompting is included.

Best for: Teams transcribing high volumes of recorded audio who don't need the top accuracy tier.

Pros

βœ… The cheapest transcription rate AssemblyAI publishes

βœ… 99-language support at the base rate

Cons

❌ No support for free-form prompting, which the Pro model does support

❌ Add-ons like medical mode cost as much as the base rate itself

Universal-3.5 Pro: $0.21/hr

What's included: AssemblyAI's flagship async model, built for 18 languages with native code switching and its most accurate speaker diarization to date. Keyterms and free-form prompting each add $0.05/hr. Diarization adds $0.02/hr, and medical mode adds $0.15/hr.

Best for: Teams that need the highest transcription accuracy for compliance-sensitive or multilingual recordings. Our guide to real-time audio processing APIs covers how AssemblyAI's streaming models compare on latency-sensitive work.

Pros

βœ… Native code switching across 18 languages

βœ… AssemblyAI's most accurate diarization model to date

Cons

❌ Narrower language coverage than Universal-2's 99 languages

❌ Every add-on (prompting, diarization, medical mode) costs extra

Universal-Streaming: $0.15/hr

What's included: AssemblyAI's cost-effective real-time model, with English and Multilingual variants for production voice applications. Keyterms prompting adds $0.04/hr, and diarization adds $0.12/hr. Free-form prompting isn’t supported on this model.

Best for: Voice applications where latency matters more than multilingual coverage.

Pros

βœ… The lowest streaming rate AssemblyAI offers

βœ… Built specifically for low-latency, production voice use

Cons

❌ The base variant is English-only, and the Multilingual variant covers six languages

❌ Diarization costs six times more here than on the pre-recorded models

Universal-3.5 Pro Realtime: $0.45/hr

What's included: AssemblyAI's most accurate real-time model, with built-in context carryover, conversation memory, and self-correcting speaker labels across 18 languages.

Keyterms prompting is included. Diarization adds $0.12/hr, Voice Focus adds $0.10/hr, medical mode adds $0.15/hr, free-form prompting adds $0.05/hr, and PII text redaction adds $0.12/hr.

Best for: Production voice agents that need top-tier real-time accuracy and can absorb the higher base rate.

Pros

βœ… Context carryover and conversation memory built in

βœ… Keyterms prompting is included, unlike the value-tier streaming model

Cons

❌ Triple the rate of Universal-Streaming

❌ Voice Focus and medical mode both add to the per-hour cost

Sync API: $0.45/hr

What's included: A newly launched single-call model that delivers Universal-3.5 Pro accuracy with a median response time around 134ms, for requests up to two minutes long. Keyterms prompting and conversation context are both included. Free-form prompting adds $0.05/hr.

Best for: Applications that need one fast, synchronous transcription call rather than a streaming connection.

Pros

βœ… Conversation context is included, which most other models charge for

βœ… Sub-150ms median response time for short requests

Cons

❌ Capped at two minutes per request

❌ Same rate as the Pro Realtime model, with a narrower use case

Dictation API: $0.62/hr

What's included: A new single-utterance model built on Universal-3.5 Pro across 19 languages, for clips up to 120 seconds.

It returns finished text with filler removed and self-corrections resolved, alongside the verbatim transcript. Keyterms prompting, prompting, and output instructions are all included.

Best for: Apps that turn spoken input into ready-to-send text, like messages, notes, or task lists.

Pros

βœ… Every feature is included in the $0.62/hr rate

βœ… Returns both cleaned text and the verbatim transcript

Cons

❌ The highest per-hour rate of any AssemblyAI transcription product

❌ Capped at 120 seconds per request

Voice Agent API: $4.50/hr ($0.075/min)

What's included: A fully managed voice agent stack that bundles Universal-3.5 Pro Realtime transcription, AssemblyAI's proprietary conversational LLM and TTS, advanced turn detection, interruption detection, and hosting into one rate.

Recordings and transcripts are included, and telephony runs on your own Twilio account with no markup on the carrier rate.

Best for: Teams that want a single managed agent platform instead of assembling STT, LLM, and TTS separately.

Pros

βœ… No per-layer add-ons, concurrency fees, or per-agent subscriptions

βœ… SIP trunking through your own Twilio account, at your existing carrier rate

Cons

❌ At $4.50/hr, this is over 20 times the base transcription rate

❌ Idle time and a 30-second reconnect window on dropped sessions are both billable

AssemblyAI Pricing for Speech Understanding and Guardrails

These models run on a transcript and bill per hour of audio, on top of the speech model rate.

Add-onPrice
Speaker Identification$0.02/hr (low effort) Β· $0.10/hr (medium effort)
Translation$0.06/hr
Custom Formatting$0.03/hr
Entity Detection$0.08/hr
Sentiment Analysis$0.02/hr
Key Phrases$0.01/hr
Topic Detection$0.15/hr
Summarization$0.02/hr (low effort) Β· $0.07/hr (medium effort)
Profanity Filtering$0.01/hr
PII Audio Redaction$0.05/hr
PII Text Redaction$0.08/hr (pre-recorded) Β· $0.12/hr (streaming)
Content Moderation$0.15/hr

Custom/Enterprise: Contact Sales

What's included: Custom rate limits, enhanced concurrency, and volume discounts tailored to high-volume workloads, across every product line.

Startups can apply to the AssemblyAI Startup Program, Y Combinator companies qualify for special pricing, and AWS Marketplace billing is available at list price, with private offers for volume rates.

Best for: Organizations with volume or compliance requirements that self-serve pricing doesn't cover.

Pros

βœ… Volume discounts that can lower per-hour costs at scale

βœ… A dedicated Forward Deployed Engineer and custom voices on Voice Agent API contracts

Cons

❌ No published rate, so you can't budget precisely without a sales conversation

❌ Requires a direct relationship with AssemblyAI's sales team

How AssemblyAI Pricing Bills Your Usage

  • Pre-recorded audio is pro-rated to the exact second, and failed transcripts are not charged.
  • Streaming bills on how long the WebSocket stays open, including idle time. A session that is never closed auto-closes after 3 hours and bills for all 3 hours.
  • Multichannel audio bills per channel, so a 1-hour stereo file bills as 2 hours.
  • Usage draws down a prepaid balance. At $0 without auto-pay, API access pauses until you top up.

Which AssemblyAI Pricing Option Should You Choose?

Choose Universal-2 if you:

  • Process high volumes of recorded audio and want the lowest per-hour rate
  • Don't need free-form prompting or the newest diarization model

Choose Universal-3.5 Pro if you:

  • Transcribe multilingual audio with frequent code switching
  • Need the highest available accuracy for compliance or medical use cases

Choose Universal-Streaming if you:

  • Build English-only, latency-sensitive voice applications
  • Want the lowest real-time rate AssemblyAI publishes

Choose the Sync API if you:

  • Need a single fast call instead of an open streaming connection
  • Work with short audio clips under two minutes

Choose the Voice Agent API if you:

  • Want a fully managed voice agent without assembling STT, LLM, and TTS separately
  • Value one flat rate over negotiating discounts across three vendors

What AssemblyAI Costs on a Real Workload

  • Recorded calls: 500 hours a month of mono call audio on Universal-3.5 Pro, with diarization ($0.02), PII text redaction ($0.08), and sentiment analysis ($0.02), costs $0.33/hr, or $165 a month.

    Send the same calls as 2-channel stereo, and the base rate alone doubles to $210.

  • Voice agent: 10,000 three-minute calls a month is 500 hours of session time. Universal-3.5 Pro Realtime with Voice Focus costs $0.55/hr, or $275 a month for STT only, before your LLM and TTS. The Voice Agent API covers all three for $2,250 a month.

Is AssemblyAI Worth the Cost?

AssemblyAI pricing is easy to read per product, but the complexity is spread across eight priced products and a long add-on menu, so the sticker rate on the homepage rarely matches the final invoice.

Per Cekura's Converse-STT benchmark, AssemblyAI Universal 3.5 Pro had the lowest word error rate of 15 English speech models on 1,000 public Pipecat voice-agent clips: 1.93%, against 3.46% for Deepgram Nova-3.

On eight licensed Ocular recordings, it scored 4.23%, ninth of 15, and its median final-text delay was 180 ms.

AssemblyAI is worth it if you:

  • Want pre-recorded audio billed to the exact second, with no contracts, minimums, or monthly subscriptions
  • Are building a voice agent and want STT, LLM, and TTS bundled at one flat rate

Skip AssemblyAI (or pair it with something else) if you:

  • Only need basic transcription without add-ons, where a flat per-minute competitor may undercut it
  • Need text-to-speech as a standalone product outside the Voice Agent API bundle

AssemblyAI Pricing vs. Deepgram and Gladia

πŸ›  ToolπŸ’° Starting Price🎯 Best For
AssemblyAI$0.15/hr (Universal-2, pre-recorded)Bundled voice agent option alongside standalone transcription
Deepgram~$0.258/hr ($0.0043/min, Nova-3 pre-recorded)One vendor across STT, TTS, and voice-agent orchestration
Gladia$0.61/hr (Starter, async)All-inclusive pricing with diarization bundled in

Deepgram's pay-as-you-go rate for Nova-3 pre-recorded is $0.0043/min. Diarization is included on pre-recorded audio but costs $0.0020/min on streaming. Our Deepgram pricing breakdown covers every Deepgram tier and add-on.

Gladia's Starter plan costs $0.61/hr for async and $0.75/hr for real-time, with diarization and every other feature included, so it starts higher than AssemblyAI's Universal-2 but has no add-on math.

For a look at how a bundled TTS-plus-agent platform prices out, see our ElevenLabs pricing breakdown.

Cekura vs. AssemblyAI: Which Should You Choose?

Cekura is better for: Independent testing and monitoring on top of whatever speech stack you run. Cekura simulates conversations and reviews production calls, whether your agent runs on AssemblyAI, Deepgram, or another provider entirely.

AssemblyAI is better for: The transcription and voice-agent layer itself. If the speech layer is the gap in your stack, AssemblyAI is the product to evaluate, not a substitute for testing it once it's live.

Our breakdown of how voice assistants process human language covers where transcription fits in that pipeline.

Use both if: you're running an AssemblyAI-powered agent in a regulated setting like healthcare, where a missed branch in a call flow means a real appointment gets missed, and manual review only covers a small slice of production calls.

Cekura tests, monitors, and improves an AssemblyAI-powered agent in three stages:

  • Testing: Cekura runs simulated callers against the agent before launch and under concurrent load, and A/B tests two agent versions on the same scenarios, so you can compare an AssemblyAI build with a Deepgram build before you commit.
  • Monitoring: Cekura scores live calls for instruction-following, sentiment, latency, and drop-off once the agent is in production.
  • Improvement: Cekura diagnoses failing calls, proposes prompt and configuration fixes, and re-tests them

Cekura works alongside any speech stack, with native integrations for LiveKit, Pipecat, Vapi, Retell, and ElevenLabs.

Cekura supports SOC 2, HIPAA, and GDPR compliance, covering transcript redaction, role-based access, and audit trails.

The Bottom Line on AssemblyAI Pricing

AssemblyAI pricing is transparent at the product level. Pre-recorded audio is pro-rated to the exact second, but streaming and the Voice Agent API bill on how long the WebSocket stays open, including idle time. The real cost variables are the add-on stack and how cleanly your app closes sessions.

Before you commit, model your actual mix of diarization, medical mode, or Voice Agent API usage against the base rate. That math moves the bill more than any single pricing decision on this page.

Once your numbers hold up, test the agent that runs on them. Book a demo to see how Cekura runs simulated callers against your AssemblyAI-powered agent, and flags missed turns, wrong answers, and dropped calls before real customers hit them.

Frequently Asked Questions

How much does AssemblyAI cost per hour?

AssemblyAI pricing starts at $0.15 per hour for Universal-2 pre-recorded transcription and $0.21 per hour for Universal-3.5 Pro.

Streaming starts at $0.15 per hour, the Dictation API costs $0.62 per hour, and the Voice Agent API costs $4.50 per hour ($0.075 per minute), before any LLM Gateway token charges.

What's included in the AssemblyAI free tier?

New AssemblyAI accounts get $50 in free credits with no credit card required, worth up to 185 hours of pre-recorded or 333 hours of streaming transcription at base rates. Credits do not expire, LLM Gateway usage is excluded, and streaming is capped at 5 new connections per minute on the free tier

Does opting out of AssemblyAI model training cost more?

No. On paid plans, AssemblyAI lets you opt out of its Model Improvement Program, set a time-to-live for audio and transcripts, and sign a Business Associate Agreement self-serve, at no additional cost.

Does the Voice Agent API include LLM and TTS costs?

Mostly. The Voice Agent API's $4.50/hr rate covers speech-to-text, AssemblyAI's Voice Agent LLM, text-to-speech, and hosting in one meter, billed per second of open session time. Any LLM Gateway calls the agent makes are billed separately at token rates.

Is AssemblyAI cheaper than Deepgram?

It depends on the workload. AssemblyAI's Universal-2 rate of $0.15/hr (about $0.0025/min) undercuts Deepgram's pre-recorded Nova-3 rate of roughly $0.0043/min.

Add-ons like diarization and medical mode change the comparison depending on which features each workload needs.

Does AssemblyAI offer volume discounts?

Yes. AssemblyAI offers custom pricing with volume discounts, enhanced concurrency, and custom rate limits for high-volume workloads through its sales team, across every product line.

Ready to ship voice
agents fast?Β 

Book a demo