CartSignal
Public-data review · Updated 2026-09-21

AssemblyAI public-data review: two meters, one rate card

AssemblyAI sells transcription by the hour with no plan and no minimum, from $0.15 to $0.45, plus a bundled voice agent at $4.50. It markets per-second billing where you pay for exactly the audio you send. That is true of one of its two products. Streaming is billed on how long the connection stays open, not on the audio going through it — idle time included — and a session nobody closes bills for three hours. The second thing no roundup carries: the model that costs 40% more covers 81 fewer languages.

Two glass flasks side by side on a dark teal surface joined by a single curved glowing pipe, the left flask marked with faint graduation lines and holding only a shallow pool of cyan liquid fed by a thin trickle, the right flask filled to the top with bright glowing cyan liquid even though the pipe above it is empty
Visit AssemblyAITranscription cost per hour

Quick verdict

Review type: Public-data review, not hands-on testing. Every figure below is read off AssemblyAI's own pricing page and developer documentation on 21 September 2026. Nothing here is a tracker figure — the pricing page renders its rates, its add-on card and its LLM token rates in full, which is not true of several vendors in this directory.

Pricing note: Usage only. No subscription, no credit pool, no seat, no annual discount and no tier that costs more per unit than the one beneath it, because there are no tiers. What replaces the plan-ladder traps this site keeps finding elsewhere is a metering difference between two products sold from the same page.

Disclosure: No affiliate relationship is recorded for this listing in the local CartSignal data. The vendor link above is a plain official URL.

The rate card

All rates are per hour, from the official pricing page today. The right-hand column is what $50 of free signup credit buys at each rate.

ModelModePer hourLanguages$50 free credit buys
Universal-2Pre-recorded$0.1599333 hours
Universal-3.5 ProPre-recorded$0.2118238 hours
Universal-Streaming EnglishStreaming$0.151333 hours
Universal-Streaming MultilingualStreaming$0.156333 hours
Universal-3.5 Pro RealtimeStreaming$0.4518111 hours
Voice Agent APIBundle$4.5011 hours

Two things are worth noting before any comparison. The transcription card spans exactly 3x from $0.15 to $0.45, which is modest — the entire self-serve field this site priced yesterday only spans 6.1x, so AssemblyAI's internal spread is half the whole market's. And the streaming premium is not one number: Universal-Streaming matches Universal-2's batch price exactly, a 0% premium, while Universal-3.5 Pro Realtime is 114.3% above its own batch equivalent. Same vendor, same page.

The headline: two meters, and the advertised one covers half the product

AssemblyAI's marketing is specific about how it bills. Its own writing on transcription at scale says "per-second billing means you don't pay rounded-up time per request", and that is a real advantage against a vendor like Google, whose documentation rounds every request up to the nearest 15 seconds. On pre-recorded audio the claim holds exactly. The billing documentation states that "Pre-recorded transcription is billed on the duration of the submitted audio or video file in seconds, multiplied by the hourly rate for the selected speech model", pro-rated "to the exact second of audio processed", with add-ons "also pro-rated to the exact second" — and, unusually generously, "Credits are deducted only after a successful transcription completes. If a request errors, you aren't charged for it."

Streaming runs on a different meter, and AssemblyAI documents it just as plainly: "Streaming Speech-to-Text is billed on the total duration that your WebSocket connection stays open, not on the amount of audio you send." The dedicated session-pricing FAQ removes any ambiguity: "You're charged for idle time on an open session just the same as you are for time when audio is actively flowing", and "Each open session is billed independently, so concurrent sessions accumulate billed time in parallel."

This is not a rounding rule, and it is not a minimum. It is a different quantity being measured. On pre-recorded audio you buy audio; on streaming you buy wall-clock time on a socket. The consequence is that the published rate is a floor you only reach if someone talks continuously for the whole session, which in a live product essentially never happens:

Share of session with audio flowingEffective multiplierUniversal-Streaming, effectiveU-3.5 Pro Realtime, effective
100%1.00x$0.150$0.450
80%1.25x$0.188$0.563
60%1.67x$0.250$0.750
50%2.00x$0.300$0.900
40%2.50x$0.375$1.125
25%4.00x$0.600$1.800

The multiplier is rate-independent, so it applies whatever the model costs. It is also the reason a dual-stream setup doubles the bill for one conversation, which AssemblyAI illustrates itself: "a single call that is dual-streamed under two separate session IDs for 5 minutes bills as 10 minutes of session time." A support call where the caller and agent each get a session is billed twice over, and that is documented behaviour rather than a bug.

The sharpest edge is the unterminated session. "Sessions that are not terminated auto-close after 3 hours, and you'll be billed for the full 3-hour session duration regardless of how much audio was actually streamed." That is $0.45 on Universal-Streaming and $1.35 on Universal-3.5 Pro Realtime for a connection that may have carried nothing. AssemblyAI names this as the usual cause of surprise bills and gives the fix in one line — "Always send a session termination message when your application is finished with a stream. This is what stops billing for the session" — but a crashed client or a dropped mobile connection does not send that message. Thirty-seven leaked Realtime sessions exhaust the entire $50 free credit.

Correction to this site, 21 September 2026

CartSignal's transcription cost page, published yesterday, said that neither AssemblyAI, Deepgram nor ElevenLabs publishes a billing increment at all, and advised readers to "assume per-second and verify on your first invoice". That is wrong for AssemblyAI, which publishes two billing rules explicitly in its documentation — and for its streaming product, assuming per-second billing of audio is precisely the wrong assumption. That page has been corrected today. The broader advice stands for the vendors where it was accurate.

Price and coverage run in opposite directions

This site found the same shape at Amazon Polly yesterday, where every step up the price ladder removed languages and regions. AssemblyAI does it in one step. The supported-languages documentation states that "Universal-3.5 Pro supports the following 18 languages" and that "Universal-2 supports 99 languages." So the newer model costs 40% more and reaches 18.2% of the languages of the one beneath it. Per language covered, that is $0.0117 against $0.0015 — a 7.7x gap.

The older model is also the more capable one on features. Both Summarization and Auto Chapters are marked on the pricing page as Universal-2 only and deprecated, so the $0.21 model cannot summarize or chapter at any price while the $0.15 model can. What the extra 40% buys is accuracy and code switching, which AssemblyAI backs with published measurements rather than adjectives: in its launch post of 7 July 2026 it reports 7.69% normalized word error rate on code-switched audio against ElevenLabs Scribe v2 at 8.77%, Deepgram Nova-3 Multilingual at 12.22% and OpenAI GPT-4o Transcribe at 44.58%, and 30.17% cpWER on diarization against Deepgram Nova-3 English at 37.92%, Scribe v2 at 35.26% and Gladia at 36.87%. These are the vendor's own numbers on the vendor's own test sets and are reproduced here as claims, not findings.

AssemblyAI papers over the coverage gap automatically: "Universal-3.5 Pro supports 18 languages, and for anything outside that set, the system automatically falls back to Universal-2, giving you coverage across 99 languages total". That is genuinely useful engineering and it leaves one question unanswered. Nothing CartSignal could find states which rate a fallback request is billed at. A Vietnamese file is $0.21; a Ukrainian one silently transcribed by Universal-2 could be $0.21 or $0.15, a 40% difference on a workload that gives no signal it has switched models. If your audio is mixed-language, price it at the higher rate and check an invoice.

The add-on card, and why you cannot switch everything on

AssemblyAI publishes the most complete add-on list in this field. Every item is quoted per hour and stacks additively on the base rate.

Add-onPer hourAdd-onPer hour
Profanity filtering+$0.01Speaker diarization, async standard+$0.02
Key phrases / auto highlights+$0.01Speaker diarization, async experimental+$0.065
Speaker identification+$0.02Speaker diarization, streaming+$0.12
Sentiment analysis+$0.02Keyterms prompting, async+$0.05
Custom formatting+$0.03Keyterms prompting, Streaming English+$0.04
Summarization (Universal-2 only, deprecated)+$0.03General prompting, beta+$0.05
PII audio redaction+$0.05Voice focus, U-3.5 Pro Realtime+$0.10
Translation (pre-recorded only)+$0.06Topic detection, IAB+$0.15
Entity detection+$0.08Content moderation+$0.15
PII text redaction+$0.08Medical mode+$0.15
Auto chapters (Universal-2 only, deprecated)+$0.08

A realistic basket moves the bill more than the model choice does. Speaker diarization, entity detection, sentiment analysis and custom formatting together are +$0.15 an hour, which is $0.36 on Universal-3.5 Pro (+71.4%) and $0.30 on Universal-2 (+100%, because the same basket is a larger fraction of a smaller base).

The "everything on" figure needs care, because several add-ons are exclusive to one model or one mode. Switching on every add-on that can actually apply produces a result worth stating: both async models land on exactly $0.99 an hour — Universal-3.5 Pro at $0.21 plus $0.78 of applicable add-ons, Universal-2 at $0.15 plus $0.84, since it gains Summarization and Auto Chapters but cannot run General Prompting. Fully loaded, the 40% base-rate gap disappears completely. With medical mode both are $1.14.

Two structural limits are easy to miss. The whole Speech Understanding group — speaker identification, translation, custom formatting, entity detection, sentiment, key phrases and topic detection — is marked pre-recorded only, so a live pipeline cannot buy any of it. And fully loading a stream is cheap by comparison precisely because so little is purchasable: Universal-Streaming English with diarization and keyterms is $0.31, Universal-3.5 Pro Realtime with everything available is $0.72.

The same feature, three prices

Speaker diarization is the clearest example of a pattern worth checking on any usage-priced vendor: the feature name is constant and the rate is not. Async standard is +$0.02, async experimental is +$0.065 (3.25x) and streaming is +$0.12 (6x). Measured against the base it sits on, diarization is 9.5% of a Universal-3.5 Pro hour and 80% of a Universal-Streaming hour, which takes a diarized live transcript to $0.27 and quietly erases most of the advantage of the cheap streaming tier.

Keyterms prompting runs the other way and produces the oddest line on the card. It costs +$0.05 on both async models and +$0.04 on Universal-Streaming English, but it is included at no charge on Universal-3.5 Pro Realtime and on Universal-Streaming Multilingual. Those last two sit at the same $0.15 base as the English model. So the English-only streaming model charges 26.7% more than its multilingual sibling for an identical feature set, while covering one language instead of six. On public data there is no configuration in which Universal-Streaming English is the better buy, and CartSignal could not find a stated reason for the difference.

The Voice Agent API is 30x its own speech recognition

AssemblyAI sells a bundled agent at $4.50 an hour ($0.075 a minute), described as unified billing across speech to text, LLM reasoning, text to speech, turn detection, interruption handling and tool calling. Measured against the company's own rate card that is 30x its cheapest streaming transcription, 21.4x its top pre-recorded rate and 10x its dearest streaming model — so the speech recognition AssemblyAI is actually known for accounts for about 3.3% of the price of the bundle built around it.

Unlike Deepgram, whose bundle this site decomposed in full because it publishes rates for every component, AssemblyAI can only be taken apart at one end: it publishes no text-to-speech rate, so the largest unknown stays unknown. It does publish LLM token rates through its gateway, spanning $0.05 per million input tokens for GPT-5 Nano to $5.00 for GPT-5.5 and Claude Opus, with a note that the quoted rates are for global routing and that in-region US or EU routing is 10% higher — a rare published price for data residency, and worth knowing before a procurement conversation assumes it is free. Bringing your own LLM is supported; the documentation references an example using "Claude through the AssemblyAI gateway, or your own endpoint".

One coincidence is worth recording because it suggests a settled market rate rather than a copied one. AssemblyAI's $0.075 agent minute is identical to Deepgram's standard managed agent minute after the increase of 15 September 2026, from a vendor setting prices independently on a different card. This site found the same convergence in transcription add-ons, where keyterm prompting is $0.05 an hour at both AssemblyAI and ElevenLabs.

Training and retention: opted in unless you say otherwise

The data retention and model training page states that "Only certain files submitted to the API, as permitted by the applicable contract, are used for model training", and that "We will not use files you submit for model training if you are subject to a Business Associate Addendum, are utilizing our European servers, or if you have opted out." The default is therefore participation, and the opt-out is a deliberate action.

It also gates the strictest retention setting behind that same choice: "If you are opted out of model training, we offer zero data retention of audio and transcripts for our Streaming product." Left alone, async audio deletion begins at 24 hours and completes within 48, and final transcripts begin deleting at 30 days. For anyone handling client recordings under an NDA, the practical reading is that the opt-out is not a privacy nicety but the switch that changes what is stored.

That is the reverse of the default this site recorded at Amazon Polly, where AWS states it does not use inputs or outputs to train the service at all, and closer in shape to Google's Gemini API, where the data terms depend on whether billing is switched on. The broader question of what you may then do with the output is covered on commercial use rights for AI voice and video.

Where AssemblyAI sits in the field

On base rates it is at the cheap end and the field is tight. Against the per-hour ranking on what AI transcription costs, Universal-2 and Universal-Streaming at $0.15 sit just above Inworld's $0.10 and below ElevenLabs Scribe v2 at $0.22, Deepgram Nova-3 streaming at $0.288 and Cartesia Ink-2 at $0.40 to $0.54. Universal-3.5 Pro Realtime at $0.45 is near the top.

What actually distinguishes it is not the rate. It is the completeness of the published card — add-ons, token rates, retention defaults and the billing rules themselves are all on pages you can read without a sales call, which is more than several vendors in this directory manage. The cost of that transparency is that the details matter: the meter changes between products, a feature costs three different amounts, and one model is dominated by its own sibling.

Limitations and what is not established

  • Which rate a language fallback bills at. Universal-3.5 Pro falls back to Universal-2 outside its 18 languages, and no page CartSignal could find states whether that request is charged $0.21 or $0.15.
  • What the Voice Agent bundle contains. The product and documentation pages describe the capabilities but do not name the default LLM or the text-to-speech voices, and AssemblyAI publishes no TTS rate, so the bundle cannot be decomposed the way Deepgram's can.
  • How the Voice Agent API is metered. Whether it follows the session-duration rule of Universal Streaming or the audio-duration rule of pre-recorded work is not stated on the pages read. Given that it is a live product, budget on session time.
  • Concurrency. AssemblyAI's own writing claims "unlimited concurrency on standard pay-as-you-go" while the pricing page caps new streams at 100 per minute on pay-as-you-go and 5 per minute on the free tier. These measure different things — simultaneous sessions against the rate of opening them — but the two statements sit on different pages and a reader could easily take the first for the whole story.
  • The "everything on" total depends on what you count. Two add-ons are Universal-2 only and deprecated, several are pre-recorded only, and diarization and keyterms are priced per model, so there is no single all-inclusive rate. The $0.99 figures above are for the sets named, and the table is printed so the arithmetic can be checked.
  • No accuracy claim is made here. No audio was transcribed. The word error rate and diarization numbers are AssemblyAI's own, measured by AssemblyAI, and vendor benchmark posts are marketing documents in every direction. With a 6.1x field and $50 of free credit, testing your own audio is close to free.

Frequently asked questions

How much does AssemblyAI cost per hour?

Pre-recorded transcription is $0.15 an hour on Universal-2 and $0.21 on Universal-3.5 Pro. Streaming is $0.15 on Universal-Streaming English and Universal-Streaming Multilingual and $0.45 on Universal-3.5 Pro Realtime. The bundled Voice Agent API is $4.50 an hour, or $0.075 a minute. Those are base rates only. More than a dozen add-ons are quoted per hour and stack additively, so a meeting transcript with speaker diarization, entity detection, sentiment analysis and custom formatting is $0.36 an hour on Universal-3.5 Pro rather than $0.21, a 71.4% increase. Switching on every add-on that applies takes either async model to exactly $0.99 an hour.

Why is AssemblyAI streaming billed differently from batch?

Because they use different meters, and AssemblyAI documents both. Pre-recorded work is billed on the duration of the submitted file, pro-rated to the exact second, and a request that errors is not charged. Streaming is billed on the total duration that the WebSocket connection stays open rather than on the audio sent through it, and the documentation states plainly that you are charged for idle time on an open session just the same as for time when audio is actively flowing. Concurrent sessions accumulate billed time in parallel, so AssemblyAI gives its own worked example of a single call dual-streamed under two session IDs for five minutes billing as ten minutes. A session that is never terminated auto-closes after three hours and bills for all three, which is $1.35 on Universal-3.5 Pro Realtime for a connection that may have carried no audio at all.

Is Universal-3.5 Pro better than Universal-2?

It is more accurate on the benchmarks AssemblyAI publishes, and it covers far fewer languages. Universal-3.5 Pro supports 18 languages with native code switching; Universal-2 supports 99. So the model that costs 40% more reaches 18.2% of the languages of the one beneath it, and Universal-2 is also the only model that can run the Summarization and Auto Chapters add-ons, both of which AssemblyAI marks Universal-2 only and deprecated. AssemblyAI handles the coverage gap by falling back automatically to Universal-2 for anything outside the 18, which is convenient but leaves a billing question its documentation does not answer: CartSignal could not establish which of the two rates a fallback request is charged at.

What does the AssemblyAI Voice Agent API include for $4.50 an hour?

AssemblyAI describes it as unified billing covering speech to text, LLM reasoning, text to speech, turn detection, interruption handling and tool calling, at $4.50 an hour or $0.075 a minute. That rate is 30 times its own cheapest streaming transcription rate and 21.4 times its top pre-recorded rate, so the speech recognition AssemblyAI actually builds is about 3.3% of the bundle it sells. The bundle cannot be fully decomposed from public data because AssemblyAI publishes no text-to-speech rate of its own, though it does publish LLM token rates through its gateway. The $0.075 minute lands on exactly the same number as Deepgram charges for its post-increase standard managed agent minute, from a vendor setting its prices independently.

Does AssemblyAI train on your audio?

By default some of it, and you have to act to stop it. The documentation states that only certain files submitted to the API, as permitted by the applicable contract, are used for model training, and that AssemblyAI will not use files you submit for model training if you are subject to a Business Associate Addendum, are utilizing our European servers, or if you have opted out. Opting out is also what unlocks the strictest retention setting, since AssemblyAI offers zero data retention of audio and transcripts for its Streaming product only to accounts that are opted out of model training. Without a Business Associate Agreement the defaults are that async audio deletion begins at 24 hours and completes within 48, while final transcripts begin deleting at 30 days. This is the opposite default from Amazon Polly, where AWS states it does not use inputs or outputs to train the service at all.

Has CartSignal tested AssemblyAI?

No. This is a public-data review built from the official AssemblyAI pricing page, its billing and pricing documentation, its Universal Streaming session-pricing FAQ, its supported-languages and Universal-3.5 Pro model pages, its data retention and model training page and its Universal-3.5 Pro launch post, all read on 21 September 2026. No audio was transcribed, so nothing here evaluates accuracy. The word error rate and diarization figures quoted are AssemblyAI measurements on AssemblyAI test sets, reproduced as vendor claims rather than independent benchmarks. Every percentage, multiple and stacked total is CartSignal arithmetic over the published rates.

Sources, all read 21 September 2026

Related CartSignal pages

What AI transcription costs per hour · Best AI voice tools by job · AI voice category · Cheapest voice, ranked per character · Deepgram · ElevenLabs · Cartesia · Amazon Polly · OpenAI audio API · Commercial use rights · AI voice cloning cost · AI-readable feed