Deepgram public-data review: the increase that skipped half the card
Deepgram sells speech by usage rather than by subscription, and its voice agent rates rose on 15 September 2026 — but only on the configurations that use Deepgram's own speech. The standard minute went $0.056 to $0.075 and advanced $0.122 to $0.163, while every bring-your-own-TTS row held its old price to the cent. That inverts the strangest line on the previous card: supplying your own text to speech used to cost 16% more than letting Deepgram do it, and now saves between 13% and 25%.
What actually changed on 15 September 2026
Updated 18 September 2026. CartSignal recorded the announced increase on 14 September, the last day of the old rates, from the promotional figures Deepgram printed beside each row. Re-reading the pricing page today shows the increase landed — but on half the card. The promotional labels and the "through 9/14" dates are gone from the voice agent rows, which now show a single plain price each.
| Voice Agent configuration | To 14 Sep (pay as you go) | Now | Change | Growth, then → now |
|---|---|---|---|---|
| Standard | $0.056/min | $0.075/min | +33.9% | $0.051 → $0.068 (+33.3%) |
| Advanced | $0.122/min | $0.163/min | +33.6% | $0.110 → $0.146 (+32.7%) |
| Standard — BYO TTS | $0.065/min | $0.065/min | unchanged | $0.051 → $0.051 (unchanged) |
| Advanced — BYO TTS | $0.122/min | $0.122/min | unchanged | $0.110 → $0.110 (unchanged) |
| Custom — BYO LLM + TTS | $0.050/min | $0.050/min | unchanged | $0.041 → $0.041 (unchanged) |
The split is not random: the two rows that rose are the two that use Deepgram's own text to speech, and every row where you supply your own held its price to the cent. An hour of standard agent conversation went from $3.36 to $4.50 on pay as you go, and an advanced hour from $7.32 to $9.78, while a bring-your-own-TTS hour still costs $3.90 and $7.32 respectively. Read as a price rise this is a third; read as a rate card it is a repricing of the speech inside the bundle, and nothing else moved.
One row could not be read reliably. Custom — BYO LLM returned three different answers across repeated fetches today — $0.059 a minute, "Contact Sales", and absent altogether — so CartSignal publishes no figure for it. Its old rate was $0.050 pay as you go and $0.041 on Growth, and the 14 September card announced $0.065 and $0.059. If you are budgeting that specific configuration, read it off the page yourself rather than from here or from any tracker; the sibling row, Custom — BYO LLM + TTS, reads consistently at $0.050 and $0.041.
Flux TTS is now billing, and carries a matching-credit offer instead
The free Flux TTS window has closed and the model now bills at $0.0450 per 1,000 characters ($0.0405 on Growth). This also retires a contradiction this page recorded on 14 September, when Deepgram's pricing page said the free window ran "until 9/14" while the Flux TTS product page and the launch post of 12 August 2026 both said "Standard pricing applies beginning September 13, 2026". The two dates no longer disagree about anything current, but if you generated Flux TTS speech on 13 or 14 September, that is the window where your invoice may not match either page, and the usage dashboard is the only place to settle it.
In its place the pricing page now carries a launch offer, quoted verbatim: "Buy Flux TTS, get matching credits. For every $1 you spend on Flux TTS, we'll add $1 in credits — up to $500. Available on Pay As You Go and Growth plans. Offer ends December 31, 2026." Taken at face value that halves the effective rate up to the cap — $500 of spend plus $500 of matched credits buys about 22.2 million characters for $500, or $22.50 per million against a $45.00 list. For the first roughly 22 million characters, in other words, Deepgram's newest and dearest voice undercuts its own mid-tier Aura-2 at $30.00 per million. Past the cap it reverts to being three times Aura-1. The offer says "credits" rather than Flux TTS credits, so whether the matched balance can be spent on transcription as well is not stated on the page.
The full published rate card
Speech to text
The streaming table is labelled "Limited-time promotional rates on streaming" and prints a regular price beside the current one. Crucially, that label carries no end date — unlike the voice agent rows, there is nothing to plan around.
| Model | Current, pay as you go | Regular, pay as you go | Per hour, current → regular |
|---|---|---|---|
| Nova-3 Monolingual, streaming | $0.0048/min | $0.0077/min | $0.288 → $0.462 |
| Nova-3 Multilingual, streaming | $0.0058/min | $0.0092/min | $0.348 → $0.552 |
| Flux English, streaming | $0.0065/min | $0.0077/min | $0.390 → $0.462 |
| Flux Multilingual, streaming | $0.0078/min | None shown | $0.468 |
| Nova-3 Monolingual, pre-recorded | $0.0043/min | Not marked promotional | $0.258 |
| Nova-3 Multilingual, pre-recorded | $0.0052/min | Not marked promotional | $0.312 |
| Whisper Large, pre-recorded | $0.0048/min | Not marked promotional | $0.288 |
The gap between Flux and Nova-3 is a discount, not a price. At regular prices, Flux English streaming and Nova-3 Monolingual streaming cost exactly the same — $0.0077 a minute on pay as you go, $0.0065 on Growth. The 35% premium Flux appears to carry today exists only because Nova-3 is discounted about 38% and Flux only about 16%. Anyone reading the current column and concluding that Deepgram's conversational turn-detecting model is inherently pricier than its general one has the wrong end of it; at list price they are the same number, and the current gap can close whenever Deepgram chooses.
Real-time gets much dearer if that promotion ends. Streaming Nova-3 Monolingual currently costs 11.6% more than the pre-recorded rate ($0.0048 against $0.0043). At the stated regular streaming price of $0.0077 it would cost 79% more than batch. If your pipeline can tolerate batch transcription, that is the difference between a modest real-time surcharge and a substantial one — and there is no published date by which to decide.
Whisper is the expensive way to buy transcription here. OpenAI's Whisper Large on Deepgram lists at $0.0048 a minute, 11.6% more than Deepgram's own newer Nova-3 Monolingual pre-recorded at $0.0043. It is also the only row on the entire card with no Growth discount at all, so for a committed customer the gap widens to 33.3% ($0.288 an hour against $0.216).
Text to speech
| Model | Pay as you go | Growth | $ per million characters | $ per minute at 1,000 chars/min |
|---|---|---|---|---|
| Aura-1 | $0.0150/1k chars | $0.0135/1k | $15.00 | $0.015 |
| Aura-2 | $0.030/1k chars | $0.027/1k | $30.00 | $0.030 |
| Flux TTS | $0.0450/1k chars | $0.0405/1k | $45.00 | $0.045 |
Deepgram's speech card is an unusually tidy 1 : 2 : 3 ladder — Aura-2 is exactly twice Aura-1, Flux TTS exactly three times it — and the Growth discount is exactly 10.0% on all three. The per-minute column uses 1,000 characters to the minute, the density ElevenLabs publishes on its own API page and the basis this site applies to every per-character vendor.
Ranked against the field in dollars per million characters, Deepgram occupies three separate rungs of one ladder: Aura-1 at $15.00 ties Inworld TTS-2 Flash near the cheap end, Aura-2 at $30.00 matches Murf's API, and Flux TTS at $45.00 sits above Cartesia Scale's $37.38 and reaches 90% of ElevenLabs' own API floor of $50.00. A 3x spread inside a single vendor's rate card is larger than the gap between many pairs of vendors, which is the general lesson this site keeps finding: the model you pick usually moves the bill more than the logo does.
That ladder is temporarily distorted at the top. Deepgram's matching-credit offer on Flux TTS, described above and running to 31 December 2026, effectively halves it to $22.50 per million characters for the first $500 of spend — which puts the newest model below Aura-2's $30.00 list rate until the cap is reached, and inverts the middle of the ladder for as long as it lasts.
The voice agent bundle against its own parts
Deepgram is one of very few vendors that publishes both a bundled agent rate and the component rates that bundle is built from, which makes the markup calculable rather than guessable. The table below sums Deepgram's own published prices for one minute of conversation: a full minute of audio transcribed with Nova-3 Monolingual streaming at $0.0048, plus the characters the agent speaks at the chosen TTS rate. Speaking density is the one assumption, so it is shown at two values rather than one.
| Components for one minute | Agent speaks 500 chars | Agent speaks 1,000 chars |
|---|---|---|
| Nova-3 streaming + Aura-1 | $0.0123 | $0.0198 |
| Nova-3 streaming + Aura-2 | $0.0198 | $0.0348 |
| Nova-3 streaming + Flux TTS | $0.0273 | $0.0498 |
| Bundled standard agent minute | $0.075 | |
| Bundle ÷ components | 6.1x to 2.7x | 3.8x to 1.5x |
The obvious objection is that the bundle also includes a language model and its own synthesis, which the component sum prices differently. The current card answers that itself, because the BYO rows are effectively a price list for the pieces. Subtracting each configuration from plain Standard values every part of the minute:
| Part of one standard agent minute | How it is derived | Pay as you go | Share of $0.075 |
|---|---|---|---|
| Speech to text | Nova-3 Monolingual streaming, list rate | $0.0048 | 6.4% |
| Text to speech | Standard − Standard BYO TTS | $0.0100 | 13.3% |
| Language model | (Standard − BYO LLM + TTS) − TTS above | $0.0150 | 20.0% |
| Orchestration and margin | the remainder | $0.0452 | 60.3% |
So about 60% of a standard agent minute is not speech and not a language model. That may well be worth paying — assembling a production agent out of raw streaming endpoints is real engineering, and turn-taking is the hard part — but it is a service fee rather than a speech cost, and it is roughly four times the speech and model combined. The same arithmetic on Growth gives $0.017 for speech and $0.010 for the model, so the split between the two moves with the plan; the subtraction is CartSignal arithmetic over Deepgram's published rows, not a breakdown Deepgram publishes.
One detail falls out of it worth noticing: the text to speech inside the advanced bundle is worth $0.041 a minute ($0.163 − $0.122), 4.1 times the cent a minute inside the standard bundle. At list rates $0.041 buys about 900 characters a minute of Flux TTS or about 1,370 of Aura-2, against roughly 670 characters of Aura-1 for the standard bundle's cent — consistent with the advanced tier shipping a more expensive voice, though Deepgram does not state which model each tier uses, so that is an inference from the prices rather than something published.
Correction: bringing your own text to speech now saves money
Updated 18 September 2026. This page previously called Standard — BYO TTS "the oddest line on the card", because at $0.065 a minute it cost 16.1% more than plain Standard at $0.056 — a penalty for removing a component Deepgram would otherwise supply. That was accurate on 14 September and is now the wrong way round. Because Standard rose to $0.075 while BYO TTS stayed at $0.065, the same two rows now read as a discount, and the same thing happened at the advanced tier and on Growth.
| Configuration | Bundled rate | BYO TTS rate | To 14 Sep | Now |
|---|---|---|---|---|
| Standard, pay as you go | $0.075 | $0.065 | 16.1% penalty | 13.3% saving |
| Standard, Growth | $0.068 | $0.051 | no difference | 25.0% saving |
| Advanced, pay as you go | $0.163 | $0.122 | no difference | 25.2% saving |
| Advanced, Growth | $0.146 | $0.110 | no difference | 24.7% saving |
Deepgram's own explainer now describes the card the way the numbers read, stating that on the Voice Agent API "swapping a component lowers the bill instead of forcing a renegotiation" and referring to "BYO LLM and BYO TTS configurations at reduced published rates". On 14 September that sentence would not have been true of BYO TTS on pay as you go. It is true today on every tier.
The practical consequence is that the increase is avoidable. If you already run a text-to-speech vendor — and plenty of teams building agents are paying ElevenLabs or Cartesia for a voice they picked deliberately — switching to the BYO TTS row leaves you on your old bill to the cent, because that row never moved. You then pay your own TTS vendor separately, which is the part the rate card does not show.
Where Deepgram sits now
The 15 September increase did not just make Deepgram dearer — it changed the ordering. Comparing like with like means comparing fully managed minutes, where the vendor supplies transcription, model and voice, because that is what Cartesia and ElevenLabs sell.
| Fully managed agent minute, no prepaid commitment | To 14 Sep | Now |
|---|---|---|
| Cartesia managed agents | $0.060 | $0.060 |
| Deepgram, standard | $0.056 | $0.075 |
| ElevenLabs Speech Engine | $0.080 | $0.080 |
Deepgram's standard minute moved from 6.7% below Cartesia to exactly 25% above it, and from 30% below ElevenLabs to 6.25% below. On that basis Cartesia now holds the cheapest fully managed agent minute listed on this site, at 20% under Deepgram. Cartesia's rate was re-verified today at "$0.06 per minute" with telephony at $0.014 a minute on a Cartesia-provided number; ElevenLabs' $0.080 is carried from the CartSignal rate card of 13 September rather than re-read today.
Two qualifications matter more than the ranking. First, telephony closes most of the gap: a Cartesia-provided number takes its minute to $0.074, within a tenth of a cent of Deepgram's $0.075, so the ordering only holds if you bring your own number. Second, Deepgram's cheaper rows are not a like-for-like rebuttal — Custom — BYO LLM + TTS at $0.050 undercuts everything here, but you then pay a language-model bill and a text-to-speech bill that none of these numbers include, so it is a lower rate for a smaller product rather than a cheaper agent. None of this is a quality judgement: these are list prices for products that differ in latency, language coverage and how much of the stack they hand you.
How you actually buy it
There are no monthly tiers. Deepgram's pricing page lists a $200 credit to start, described as carrying "No minimums. No expiration. No credit card required" — the no-expiration part is unusual enough to be worth stating, since credits that expire are the norm across this directory. The Growth plan is listed from $4,000 a year in prepaid credits "redeemed against actual usage", in exchange for the lower Growth column. Enterprise is a custom quote with cloud, private-cloud and on-premises deployment; Deepgram's documentation notes that Flux "is now available for self-hosted deployments".
What the free $200 buys, at the rates above:
| Spent on | What $200 covers |
|---|---|
| Pre-recorded transcription, Nova-3 Monolingual | ~775 hours of audio |
| Text to speech, Aura-1 | ~13.3 million characters |
| Text to speech, Flux TTS | ~4.4 million characters |
| Standard voice agent, at $0.075/min | ~44 hours of conversation |
That is a genuinely generous evaluation budget for everything except live agents, where 44 hours goes quickly under load testing. Published concurrency caps on the starting plan are up to 50 concurrent requests for speech-to-text REST, up to 45 across text-to-speech REST and websocket, and up to 45 websocket connections for Voice Agent.
Correction, 18 September 2026 — the "up to 20%" claim now checks out, on exactly one row. This page previously reported that Growth is advertised as "Save up to 20%" but that nothing on the published card reached 20%. That was true of the 14 September card and is no longer true, and the reason is the same split described above: Standard — BYO TTS held $0.051 on Growth while plain Standard rose to $0.068, so that row now carries a 21.5% Growth discount — the only row anywhere on the card that beats the headline claim, and it got there by not moving.
| Growth discount, measured row by row | Discount |
|---|---|
| Voice agent, standard — BYO TTS | 21.5% |
| Voice agent, custom — BYO LLM + TTS | 18.0% |
| Nova-3 Multilingual, pre-recorded | 17.3% |
| Nova-3 Monolingual, pre-recorded | 16.3% |
| Streaming speech to text | 12.3–13.8% |
| Aura-1, Aura-2 and Flux TTS | 10.0% each |
| Voice agent, standard and advanced | 9.3–10.4% |
| Whisper Large | 0% |
So the claim is now literally defensible — "up to 20%" is satisfied by a 21.5% row — while remaining a poor guide to what a commitment is worth. Everything a typical buyer actually spends on still sits at 9–17%, the bundled agent rows that rose are among the least discounted on the card, and Whisper Large still carries no Growth discount at all. A buyer sizing a $4,000 commitment from the table should plan on 10–17% unless their workload happens to be concentrated in that one BYO TTS row.
What Deepgram actually sells
- Nova-3 speech to text, monolingual and multilingual, in streaming and pre-recorded flavours, plus a hosted Whisper Large option.
- Flux speech to text, built for conversation, with configurable end-of-turn detection. Deepgram's docs describe tunable
eot_threshold,eager_eot_thresholdandeot_timeout_mssettings, mid-stream configuration updates and a force-end-turn signal. - Aura-1 and Aura-2 text to speech, Aura-2 described by Deepgram as "an enterprise-grade text-to-speech model built for professional, cost-effective voice at scale".
- Flux TTS, released 12 August 2026 for live conversation, with named voices including Haley, Heather, Priya, Jack, Bruce, Rufus and Drew across American, Indian and British accents, and "one conversational speech model across ten languages".
- Voice Agent API, bundling transcription, a language model and synthesis behind one per-minute meter, with bring-your-own-LLM and bring-your-own-TTS configurations.
- Cloud, private cloud and on-premises deployment, which is a genuine differentiator against most of the hosted-only vendors in this directory.
Deepgram's performance claims for Flux TTS — responses "in as low as 80ms" even under production load, a 73.4% win rate on naturalness and 77.5% on expressiveness in blind comparison against twelve competing models, and a 9.8% word error rate on hard prompts such as drug names and tracking numbers — are reproduced here as vendor claims. CartSignal has generated no audio and verified none of them. Voice cloning is not offered on the Flux TTS product page; launch coverage described it as on the roadmap alongside additional languages and finer emotional controls.
Best-fit use cases
- High-volume transcription, where $0.258 an hour pre-recorded is among the cheapest published rates in this directory
- Production voice agents where you want one meter and one vendor rather than stitching three together
- Teams that already run their own text to speech or language model, where the BYO rows did not rise on 15 September and now save 13–25%
- Regulated or air-gapped deployments needing on-premises or private-cloud speech
- Cheap programmatic narration on Aura-1, provided realism is not the deciding factor
- Evaluation on a real budget, given a $200 credit with no card and no expiry
Limitations to verify
- Streaming speech-to-text rates are still promotional with no published end date; regular prices are up to 60% higher on Nova-3, and this is now the only dated-but-undated promotion left on the card
- The Custom — BYO LLM row could not be read consistently on 18 September 2026 (it returned $0.059, "Contact Sales" and nothing across repeated fetches), so budget that configuration from the page itself
- The advertised "up to 20%" Growth saving is met on one row only (BYO TTS at 21.5%); most rows deliver 9–17% and Whisper Large delivers 0%
- The agent rates rose 33% once already, on a promotion whose end date the page did publish — the streaming discount has no such warning attached
- The Growth discount requires a prepaid commitment listed from $4,000 a year
- No voice cloning on the current Flux TTS product page
- Deepgram publishes no statement located during this review about machine-readable watermarking of generated audio, which matters for the duties in our EU AI Act labeling guide
How it fits against other tools here
Deepgram and Cartesia are the closest pair in this directory: both are developer-first, both sell speech generation and transcription plus a separately metered agent product, and both compete on latency rather than voice realism. The difference is how they package it — Cartesia runs one credit pool with agents billed separately in dollars, Deepgram runs pure usage billing with no pool at all, which makes Deepgram easier to forecast and Cartesia cheaper to start. Since 15 September, Cartesia also holds the cheaper fully managed agent minute — $0.060 against $0.075 — having changed nothing about its own price.
Against ElevenLabs, the comparison is a category apart. ElevenLabs competes on voice realism, dubbing and cloning depth and prices accordingly; Deepgram publishes no cloning and no dubbing rate at all, and its most expensive speech model still undercuts ElevenLabs' cheapest API model. If you are choosing a narration voice, that is the wrong comparison to make on price. Murf AI is the pick if you want a timeline editor rather than an API, and Speechify's developer API remains the lowest published per-character rate on this site.
The wider rankings are best AI voice tools, priced by job and the ElevenLabs alternatives shortlist, and the full category sits on the AI voice tools page.
FAQ
How much does Deepgram cost?
Deepgram sells usage, not subscriptions. As of 18 September 2026, on pay as you go: pre-recorded transcription from $0.0043 a minute ($0.258 an hour) on Nova-3 Monolingual; streaming at a promotional $0.0048 against a stated regular $0.0077; text to speech at $15.00, $30.00 and $45.00 per million characters on Aura-1, Aura-2 and Flux TTS; and a bundled voice agent at $0.075 a minute standard ($4.50 an hour) or $0.163 advanced, with bring-your-own-TTS at $0.065 and $0.122 and bring-your-own-LLM-plus-TTS at $0.050. New accounts start with a $200 credit the pricing page says has no minimums, no expiration and needs no credit card. Growth is listed from $4,000 a year in prepaid credits.
What changed in Deepgram's pricing on 15 September 2026?
The announced increase landed, but only on the rows that use Deepgram's own voice. Standard went $0.056 → $0.075 (+33.9%) and advanced $0.122 → $0.163 (+33.6%) on pay as you go, and $0.051 → $0.068 and $0.110 → $0.146 on Growth. Every bring-your-own-TTS row held its price to the cent — Standard BYO TTS is still $0.065, Advanced BYO TTS still $0.122, Custom BYO LLM + TTS still $0.050. Flux TTS stopped being free and now bills at $0.0450 per 1,000 characters, with a matching-credit offer running to 31 December 2026. The promotional labels have been removed from the agent rows, so these are plain rates now.
Does bringing your own TTS to a Deepgram agent save money?
Yes, and it did not before. Until 14 September 2026, Standard — BYO TTS at $0.065 cost 16.1% more than plain Standard at $0.056, which this page called the oddest line on the card. Because Standard rose to $0.075 and BYO TTS did not move, the same pair is now a 13.3% saving on pay as you go, 25.0% on Growth, and about 25% at the advanced tier either way. The catch is that the rate card does not show your own text-to-speech bill, which you pay separately — so it is a lower Deepgram rate for a smaller Deepgram product, not automatically a cheaper agent.
Is Deepgram's Voice Agent API cheaper than assembling the same call from Deepgram's own parts?
No. A minute transcribed with Nova-3 Monolingual streaming ($0.0048) plus 500 characters of Aura-1 speech ($0.0075) is $0.0123 in components, against $0.075 for the bundled standard minute — about 6.1x. The BYO rows let you price the rest: the bundled speech is worth $0.010 a minute and the language model about $0.015, so roughly $0.045 a minute, about 60%, is orchestration and margin rather than speech or model. Even at 1,000 characters a minute on the priciest Flux TTS voice, the bundle is still about 1.5x its own components.
Is Deepgram's Flux TTS worth three times Aura-1?
Only for live conversation. The card is a clean 1:2:3 ladder at $15.00, $30.00 and $45.00 per million characters, and Flux TTS is positioned on latency and interruption handling — Deepgram claims responses in as low as 80ms across ten languages, and first place on naturalness in its own blind comparison against twelve models, all vendor claims reproduced rather than verified. For narration rendered once and played back, none of that justifies 3x, and Aura-1 at $15.00 per million is among the cheapest rates listed on this site.
Has CartSignal tested Deepgram?
No. This is a public-data review built from Deepgram's official pricing page, product pages, developer documentation and launch post, with the full rate card re-read on 18 September 2026. No audio was generated or transcribed, and no accuracy, latency or voice quality claim is verified. Per-hour conversions, percentage changes, the component-versus-bundle arithmetic and the decomposition of the agent minute are CartSignal's, applied to Deepgram's own published numbers.
Source links
- Official Deepgram pricing page (all rates, promotional and regular prices, plan tiers, concurrency) — re-read 18 September 2026
- Official Deepgram explainer on voice agent cost ("swapping a component lowers the bill", BYO configurations "at reduced published rates")
- Cartesia pricing page (managed agent minute and telephony surcharge) — checked 18 September 2026
- Official Flux TTS product page (promotional window, voices, latency and win-rate claims)
- Official Flux TTS launch post, 12 August 2026 (free window and standard-pricing date)
- Official text-to-speech product page (Aura-2 positioning, deployment options, $0.045 per 1,000 characters)
- Official Flux documentation (end-of-turn parameters, self-hosted availability)