Best AI voice tools for creators and small businesses
"AI voice" is four different products on four different meters: narration from a script, dubbing an existing recording, a live voice agent, and transcription. Priced per minute of audio, narration runs about $0.007 to $0.20, dubbing $0.16 to $2.20, and agents $0.05 to $0.16. The bigger surprise is that the door you buy through moves the bill more than the vendor does — ElevenLabs charges about 15x more for transcription bought with subscription credits than with its own API.
Four jobs, four meters
Most "best AI voice tool" lists rank products that do not do the same job. A tool that reads your script aloud, a tool that re-voices your video in Spanish, a tool that answers your phone, and a tool that turns a meeting into text are four purchases, and every vendor prices them on a different unit.
| Job | What you give it | Unit vendors bill in | Cost range |
|---|---|---|---|
| Narration (text to speech) | A script | Characters | $0.007 – $0.20 / min |
| Dubbing (localisation) | An existing recording | Minutes of source | $0.16 – $2.20 / min |
| Voice agent (live) | A conversation | Minutes of call | $0.05 – $0.16 / min |
| Transcription (speech to text) | Audio | Hours | $0.10 – $1.02 / hour |
Read down the "unit" column and the mismatch is obvious: a price quoted per character cannot be compared to one quoted per minute of call time without a conversion, and the conversion is where most published comparisons go wrong.
First, settle what "a minute" means
Text to speech is billed per character, so turning it into a per-minute figure requires assuming a speaking density — and vendors do not agree on one.
ElevenLabs' API pricing page prints the equivalence itself: 10,000 characters is listed against "~10 min", and 20,000 against "~20 min". That is 1,000 characters to the minute, from the vendor rather than from a reviewer. Cartesia's plan allowances imply a different number — 100,000 credits against 133 advertised minutes of Sonic works out at 750 credits per minute, and the same ratio holds on all four of its tiers.
So the identical script is billed as a third more minutes at one vendor than the other. This page uses ElevenLabs' published 1,000 characters a minute as a single disclosed basis applied to every per-character vendor, so the rows are comparable to each other even where they differ from a vendor's own advertised minute count. For a density-free ranking that avoids the problem entirely, use the per-million-character table on ElevenLabs alternatives.
The finding: one vendor, two doors, same job
ElevenLabs is the only vendor here that publishes a complete credit rate card and a complete dollar rate card for the same list of products. Setting them side by side prices the packaging itself.
Its pricing page gives the credit rates; its API page gives the dollars. Pro is $99 a month for 600,000 credits, so a credit is worth $0.000165. Converting every credit rate at that value:
| Job | Subscription rate | Cost at Pro credit value | ElevenLabs API | Subscription ÷ API |
|---|---|---|---|---|
| Dubbing, automatic, watermarked | 2,000 credits/min | $0.330 / min | $0.33 / min | 1.00x |
| Dubbing, automatic, no watermark | 3,000 credits/min | $0.495 / min | $0.50 / min | 0.99x |
| Eleven Music | 900 credits/min | $0.149 / min | $0.15 / min | 0.99x |
| Text to speech, v3 / Multilingual v2 | 1 credit/char | $0.165 / 1k chars | $0.10 / 1k chars | 1.65x |
| Text to speech, Flash / Turbo | 0.5 credit/char | $0.083 / 1k chars | $0.05 / 1k chars | 1.65x |
| Voice Changer / Voice Isolator | 1,000 credits/min | $0.165 / min | $0.12 / min | 1.38x |
| Speech to text | 330 credits/min | $3.27 / hour | $0.22 / hour (Scribe v2) | ~15x |
The top three rows are the ones that make the rest trustworthy. Dubbing and music come out to the cent on both pages, which fixes $0.000165 as the reference value of an ElevenLabs credit rather than a number CartSignal chose. Measured against that same anchor, text to speech carries a consistent 1.65x subscription premium, and speech to text carries roughly a fifteenfold one.
Two honest caveats. A subscription is not a pure metered purchase — it also buys the studio interface, the voice library, the commercial terms and the included headroom, and the credits are already paid for whether you use them or not. And the higher dubbing tiers do not map cleanly between the two pages, so no equivalence is claimed for them. But the practical rule is hard to argue with: if you transcribe more than occasionally, do not spend subscription credits on it. An hour of audio costs 330 × 60 = 19,800 credits, which is two thirds of an entire Starter month for one hour of transcript.
Job 1 — Narration, where the API beats the studio
Per-character rates converted at 1,000 characters a minute, cheapest first:
| Vendor and model | $ per million characters | $ per minute |
|---|---|---|
| Amazon Polly Standard | $4.00 | $0.004 |
| Speechify developer API | $6.67 | $0.007 |
OpenAI gpt-4o-mini-tts | ~$14.40 (derived) | $0.0144 |
| Inworld TTS-2 Flash, on-demand | $15.00 | $0.015 |
| Deepgram Aura-1 | $15.00 | $0.015 |
OpenAI tts-1 | $15.00 | $0.015 |
| Amazon Polly Neural | $16.00 | $0.016 |
| Inworld TTS-2, on-demand | $25.00 | $0.025 |
OpenAI tts-1-hd | $30.00 | $0.030 |
| Deepgram Aura-2 | $30.00 | $0.030 |
| Amazon Polly Generative | $30.00 | $0.030 |
| Murf API, pay as you go | $30.00 | $0.030 |
| Cartesia Scale | $37.38 | $0.037 |
| Deepgram Flux TTS | $45.00 | $0.045 |
| Cartesia Pro | $50.00 | $0.050 |
| ElevenLabs API, Flash / Turbo / v3 Conversational | $50.00 | $0.05 |
| ElevenLabs API, v3 / Multilingual v2 | $100.00 | $0.10 |
| Murf Studio Creator (billed in time, not characters) | n/a | $0.158 |
| ElevenLabs Creator subscription | $181.82 | $0.182 |
| ElevenLabs Starter subscription | $200.00 | $0.200 |
Added 20 September 2026: the four Amazon Polly engines were missing from this table and one of them takes the bottom of it. Polly Standard at $4.00 per million characters is 40% below the previous cheapest row, which widens the spread here from 30x to 50x. Two qualifications travel with that rate and are set out on the Polly profile: thirteen of Polly's 42 language variants have no Standard voice at all, so their cheapest rate is $16.00; and Polly is billed again for word-level timing, because speech marks are returned instead of audio and charged as a second pass, which doubles the effective rate to $8.00 wherever you need subtitles or lip sync. Polly Long-form, at $100.00 per million and unavailable with speech marks, is not on this table because it is six voices in two languages.
The spread is about 50x, and the shape of it is the useful part: every developer API on the list undercuts every studio subscription on it, including within the same company. ElevenLabs' own API at $50.00 per million is four times cheaper than its own Starter plan at $200.00. The two ElevenLabs API rows are not CartSignal arithmetic — ElevenLabs prints "$0.10/minute" and "$0.05/minute" beside those tiers itself.
That does not make the subscriptions bad value. Murf Studio at $0.158 a minute is billed in Voice Generation Time because you are buying a timeline editor with word-level sync and team review, not raw synthesis; its own API is $0.030. Pay the studio premium when the interface is the product, and use the API when it is not.
One nuance worth checking on your own account: ElevenLabs' pricing page states text to speech costs "1 credit per character", with discounted models "between 0.5 and 1 credit per character". At the 0.5 end, the subscription rows above halve for Flash and Turbo — so confirm which rate your model draws before budgeting.
Added 19 September 2026 — the OpenAI row that was previously uncostable. Earlier versions of this page listed gpt-4o-mini-tts without a figure, because it bills per audio output token and no OpenAI page converts tokens to characters. OpenAI does publish the missing conversion, in its realtime cost guide rather than on the pricing page: assistant audio is "1 token per 50ms", so a minute is 1,200 tokens and $12.00 per million works out at $0.0144 a minute — 4.0% under OpenAI's own tts-1 at $15.00 per million. Treat the per-character column for that row as doubly derived; the two tts-1 rows need no assumption. Full working on the OpenAI audio API review.
Job 2 — Dubbing, billed per minute of source
Dubbing is metered on the recording going in, so a 40-minute video costs 40 minutes regardless of how much of it you keep — and it is billed again for every target language. The table below prices the audio tools only. If your source is a talking head and you need the mouth to match the new language, the video tools undercut most of this list: the full cross-vertical ranking, including HeyGen, Synthesia and the dubbing specialist Rask AI, is on what AI dubbing costs per minute.
| Option | $ per minute of source |
|---|---|
| Speechify Studio Creator (3 credits/sec) | $0.156 |
| Speechify Studio Starter | $0.208 |
| ElevenLabs Dubbing v1, watermarked (API or credits) | $0.33 |
| ElevenLabs Dubbing v1, no watermark (API or credits) | $0.50 |
| ElevenLabs Dubbing Studio, no watermark, Pro credits | $1.65 |
| ElevenLabs Dubbing v2, API | $2.20 |
| Murf Dub, Automated (2 credits/min/language at $1/credit) | $2.00 |
Two things to take from this. Removing the watermark costs exactly 50% more at ElevenLabs — the credit rates are 2,000 and 3,000 a minute — which is a real line item if you publish commercially. And dubbing spans about 14x from the cheapest option to the most expensive, so it is the job where picking the wrong tier hurts most — a one-hour course localised into one language is $9.38 through Speechify Studio Creator and $132.00 through ElevenLabs Dubbing v2.
Updated 15 September 2026: Murf Dub is now costable and has been added to the table. This page previously recorded it as "not costable — $1 per credit, no published credit-to-minute conversion"; Murf's dubbing credits article supplies the missing conversion, stating that "2 credits are consumed for every minute of your file for each language selected", which at the published $1 pay-as-you-go credit price is $2.00 per dubbed minute per language. The same article notes durations are rounded up to the nearest 30 seconds. The watermark figure above was also restated: 52% was a rounding artefact of comparing $0.33 with $0.50 when the underlying credit rates differ by exactly half.
Job 3 — Live voice agents, billed per minute of conversation
An agent minute bundles transcription, a language model and speech synthesis into one meter, which is why it costs several times a narration minute from the same vendor.
| Vendor and tier | Who supplies the model and voice | To 14 Sep 2026 | Now |
|---|---|---|---|
| Deepgram Voice Agent, BYO LLM + TTS, Growth | you supply both | $0.041 | $0.041 |
| Deepgram Voice Agent, BYO LLM + TTS, pay as you go | you supply both | $0.050 | $0.050 |
| Deepgram Voice Agent, standard BYO TTS, Growth | you supply the voice | $0.051 | $0.051 |
| Cartesia managed agents | fully managed | $0.060 | $0.060, plus $0.014 telephony on a Cartesia number |
| Deepgram Voice Agent, standard BYO TTS, pay as you go | you supply the voice | $0.065 | $0.065 |
| Deepgram Voice Agent, standard, Growth | fully managed | $0.051 | $0.068 (+33.3%) |
| Deepgram Voice Agent, standard, pay as you go | fully managed | $0.056 | $0.075 (+33.9%) |
| AssemblyAI Voice Agent API | fully managed | not checked | $0.075 (published as $4.50/hr) |
| ElevenLabs Speech Engine | fully managed | $0.080 | $0.080 at every tier |
| Deepgram Voice Agent, advanced BYO TTS, pay as you go | you supply the voice | $0.122 | $0.122 |
| Deepgram Voice Agent, advanced, pay as you go | fully managed | $0.122 | $0.163 (+33.6%) |
Read the middle column before the price columns. Only the fully managed rows are comparable with each other, because a bring-your-own row leaves you paying a separate model or voice bill that no figure here includes. Compared like with like, Cartesia at $0.060 is now the cheapest fully managed agent minute on this page, 20% under Deepgram's $0.075 and 25% under ElevenLabs — and it got there without changing its own price. Deepgram's Custom — BYO LLM row is deliberately absent: it returned three different answers across repeated reads on 18 September 2026, so it is not published here.
Added 20 September 2026: AssemblyAI publishes its Voice Agent API at "$4.50/hr ($0.075/min)", described as all-inclusive of transcription, language model, speech and orchestration — which lands it on exactly the same number as Deepgram's post-increase standard managed minute, from a vendor that sets its prices independently. It is also 21.4x AssemblyAI's own $0.21 transcription rate, the same bundle-versus-parts pattern the Deepgram profile decomposes in detail. AssemblyAI has no profile on this site yet; its transcription rates and add-on card are costed in what AI transcription costs per hour.
The option that has no per-minute rate at all
Added 19 September 2026. OpenAI's Realtime API is missing from the table above on purpose, because it cannot honestly be given a cell. It bills audio by the token, and OpenAI's cost guide states that the whole conversation is re-sent as input on every turn — "turns later in the session will be more expensive" — so the rate is a function of how long the call runs and whether prompt caching holds.
| Model and scenario | What you get | $ per minute |
|---|---|---|
OpenAI gpt-realtime-2.1-mini, first minute | model only | $0.015 |
OpenAI gpt-realtime-2.1, first minute | model only | $0.048 |
OpenAI gpt-realtime-2.1, 10-minute call, history cached | model only | $0.050 |
OpenAI gpt-realtime-2.1, 10-minute call, no cache | model only | $0.178 |
On a ten-minute call split evenly between caller and agent, the tenth minute costs 6.4x the first, and the cached and uncached totals for the identical conversation differ by 3.58x — caching is automatic but best-effort. So OpenAI is simultaneously the cheapest first minute here, 20% under Cartesia, and potentially about 3x Cartesia over a real conversation. It is also a model rather than a platform: telephony, turn handling and orchestration are yours to build, which the managed rows include. Full arithmetic and the assumptions behind it are on the OpenAI audio API review.
Updated 18 September 2026 — the increase landed on half the card. Deepgram's promotional agent rates ended on 15 September, and re-reading its pricing page today confirms standard rose 33.9% to $0.075 and advanced 33.6% to $0.163 on pay as you go. But every bring-your-own-TTS row held its old price to the cent — the two rows that rose are the two that use Deepgram's own voice. That inverts a quirk this site reported in September: supplying your own text to speech used to cost 16.1% more than letting Deepgram do it, and now saves 13.3% on pay as you go and 25.0% on Growth. Flux TTS also stopped being free at $0.0450 per 1,000 characters, though a matching-credit offer running to 31 December 2026 halves that to about $22.50 per million for the first $500 of spend. The ordering consequence for this table: Deepgram's standard minute moved from 6.7% below Cartesia's $0.060 to exactly 25% above it. Full rate card and the arithmetic are on the Deepgram public-data review.
Note also what ElevenLabs does not do here: the Speech Engine stays at $0.080 a minute on Starter, Creator, Pro, Scale and Business alike. Higher tiers buy included minutes and concurrency — 4 concurrent calls on Starter up to 30 on Business — but never a better unit rate. Agent pricing is where plan shopping saves the least.
Job 4 — Transcription, the job to keep off your credits
| Vendor and model | $ per hour of audio |
|---|---|
| Inworld STT 1, Creator tier and above | $0.10 |
| Inworld STT 1, on-demand | $0.15 |
OpenAI gpt-4o-mini-transcribe (removed 26 Feb 2027) | $0.18 |
| ElevenLabs Scribe v2, API | $0.22 |
| Deepgram Nova-3 monolingual, pre-recorded | $0.258 |
OpenAI gpt-transcribe, files | $0.27 |
| Deepgram Nova-3 monolingual, streaming | $0.288 |
OpenAI whisper-1 (removed 26 Feb 2027) | $0.36 |
| ElevenLabs Scribe v2 Realtime, API | $0.39 |
| Deepgram Flux English, streaming | $0.39 |
| Cartesia ink-2, Pro plan credits | $0.54 |
OpenAI gpt-live-transcribe, streaming | $1.02 |
| ElevenLabs subscription credits, Pro | $3.27 |
| ElevenLabs subscription credits, Starter | $3.96 |
Every developer API here lands between $0.10 and $1.02 an hour — a 10.2x spread, and cheap enough at the bottom that transcription is close to a rounding error for most creators. Buy the same hour with ElevenLabs subscription credits and it is $3.27, about 33x the cheapest API rate. Added 19 September 2026: the four OpenAI rows widen the spread and sit at both ends of it. Batch file transcription on gpt-transcribe at $0.27 is mid-field, but streaming on gpt-live-transcribe at $1.02 is 2.6x the $0.39 streaming rates above it and nearly double the dearest other row on the table — so if live captions are the product, OpenAI is not the cheap option. Two of its rows also expire: whisper-1 and gpt-4o-mini-transcribe are scheduled for removal on 26 February 2027, and whisper-1 is the only OpenAI model documented to return word or segment timestamps at all, which matters more than its price if you generate subtitles. Deepgram's streaming figures are also marked "current price" against higher regular prices ($0.0077 a minute monolingual, or $0.462 an hour), so treat those two rows as promotional — and unlike its agent rates, that promotion carries no published end date. Two consequences worth reading off the regular column: real-time Nova-3 would go from an 11.6% premium over the batch rate to 79%, and Flux English and Nova-3 Monolingual carry the identical regular price of $0.0077, so the gap between those two rows above is a difference in discount depth rather than in list price. Detail on the Deepgram public-data review.
Added 20 September 2026: this table ranks the vendors CartSignal already profiles, and a wider sweep of the transcription market changes three things about it. First, two vendors undercut or match the bottom of it — AssemblyAI at $0.15 an hour on Universal-2 and $0.21 on Universal-3.5 Pro, now profiled here, and Gladia at $0.61 with speaker diarization already in the price. Second, the streaming premium is not a constant: AssemblyAI charges the same $0.15 for its Universal-Streaming models as for batch, against 11.6% at Deepgram, 77.3% at ElevenLabs and 278% at OpenAI. Third, and the one most likely to affect a real invoice, the base rate is not the price — AssemblyAI publishes more than a dozen per-hour add-ons, so a meeting transcript with diarization, entity detection, sentiment and custom formatting costs $0.36 rather than $0.21. And a fourth, which applies to the streaming rows specifically: AssemblyAI bills a live session on how long the connection stays open rather than on the audio sent through it, idle time included, so its $0.15 streaming rate doubles on a conversation that carries speech half the time and a session nobody closes bills for three hours — set out in the AssemblyAI profile. Full ranking, the add-on arithmetic, the medical premiums that span 0% to 1,150%, and the 15-second rounding rule that can bill a three-second clip as fifteen: what AI transcription costs per hour.
Picks by job
- YouTube and podcast narration, quality first: ElevenLabs — but call it through the API at $0.05–$0.10 a minute rather than buying Starter at $0.20, unless you want the studio and voice library.
- Narration inside a timeline, with team review: Murf Studio. You are paying $0.158 a minute for the editor, not the synthesis.
- High-volume programmatic narration: Speechify's developer API is the lowest published rate on this page at about $6.67 per million characters.
- Low-latency app and agent voice: Cartesia at $0.06 a minute, which since 15 September 2026 is the cheapest fully managed agent minute here. Deepgram undercuts it only in configurations where you bring your own language model or voice — and if you already run a text-to-speech vendor, Deepgram's BYO TTS row at $0.065 never took the September increase at all.
- Localisation on a budget: Speechify Studio at about $0.156 a minute of source, well under ElevenLabs' $0.33 floor.
- Transcription: any developer API. Inworld at $0.10 an hour and ElevenLabs Scribe v2 at $0.22 are both far below what a subscription charges.
Two listings that are no longer buyable
Older AI-voice roundups still recommend both of these, so they are kept here as corrections. PlayHT is gone — Meta acquired the PlayAI team in July 2025, the TTS service was deprecated within weeks, and play.ht no longer resolves. Resemble AI announced on 13 August 2026 that it is "not selling voice AI to new customers" and moved the company to deepfake detection and watermarking; its voice models remain free and self-hostable under the MIT licence as Chatterbox. What to use instead covers the migration.
A third option worth knowing about but not costed here: Mistral's Voxtral TTS, released with open weights in 2026. Mistral's pricing page states that speech models are billed per minute but published no readable rate when checked on 13 September 2026, so no figure appears above — and open weights do not automatically mean free commercial use, so check the licence attached to the weights before shipping.
Before you publish: consent and labelling
Two rules bind AI voice more tightly than AI video. Cloning a voice requires the speaker's consent, and vendors are tightening how they verify it — Speechify replaced a self-attestation checkbox with a challenge-response recording flow in August 2026. And since 2 August 2026, EU AI Act Article 50 requires generative audio to be marked machine-readably. ElevenLabs began SynthID audio watermarking in June 2026; several vendors above publish no clear statement. Note too that ElevenLabs prices watermark removal as a 50% dubbing surcharge, which is a compliance decision as much as a budget one. Full detail: labelling AI video and voice under the EU AI Act.
Consent is not only a compliance question — on two vendors it is a hard product limit. ElevenLabs states in its cloning documentation that "You can only create a Professional Voice Clone of your own voice. Even with their consent, you cannot clone someone else's voice", and Synthesia likewise allows only your own. No plan removes either restriction, so if the voice belongs to a client or a hired actor those two are out before price enters the decision. Speechify is built the other way, issuing a single-use phrase for the speaker to read aloud as the consent record, with no cap on how many voices you create. Which plans include a clone slot, what each vendor demands as proof, and how much audio you must supply: what AI voice cloning costs.
A third rule binds before either of those, and it is contractual rather than legal: most free tiers here do not grant you commercial use of what you generate. ElevenLabs states that a Free User "may only use the Services for non-commercial purposes", and its help centre adds that content made outside a paid subscription "(before or after)" never becomes commercial — so a free evaluation library has to be regenerated once you pay. Murf gates it in the product instead by withholding the export. Clause by clause: can you use AI voice and video commercially?
Test script before paying
Use one 45–60 second script across every candidate. Include a product name, one tricky pronunciation, one number read aloud, one acronym, and one sentence with an em-dash pause. Generate it on each tool, then check pronunciation, pacing, breath, and how much of your credit allowance a single minute consumed. For agents, run one three-turn interruption test — latency and barge-in behaviour matter more than voice quality on a phone call.
Frequently asked questions
What does one minute of AI voice actually cost?
It depends which of four jobs you are buying. Narration from a script runs about $0.007 a minute on the Speechify developer API up to $0.20 on an ElevenLabs Starter subscription. Dubbing runs $0.16 to $2.20 a minute of source. A live voice agent runs $0.05 to $0.16 a minute of conversation. Transcription runs $0.10 to $1.02 an hour on a developer API.
Is an ElevenLabs subscription more expensive than the ElevenLabs API?
For some jobs, by a lot. Converted at Pro's $0.000165 credit value, the two rate cards match to the cent on dubbing and music, run 1.65x higher by subscription for text to speech, and about 15x higher by subscription for speech to text — $3.27 an hour against $0.22 on Scribe v2. The subscription also buys the studio, the voice library and the commercial terms, but transcription in particular should not come out of credits.
How many characters are in a minute of speech?
Vendors disagree, which is why per-minute comparisons between them mislead. ElevenLabs' API page equates 10,000 characters with about 10 minutes — 1,000 to the minute. Cartesia's allowances imply 750. The same script is billed as a third more minutes at one vendor than the other, so compare per million characters, or per minute on one disclosed basis applied to everyone.
Which AI voice tool is cheapest?
No tool wins all four jobs. Cheapest narration is a developer API rather than any studio subscription; cheapest costable dubbing is Speechify Studio at about $0.156 a minute; cheapest transcription is Inworld at $0.10 an hour. For agents the answer changed on 15 September 2026, when Deepgram's promotional rates ended and its standard minute rose 33.9% to $0.075: the cheapest fully managed agent minute is now Cartesia at $0.060, 20% under Deepgram and 25% under the ElevenLabs Speech Engine at $0.080. Deepgram publishes cheaper rows — $0.050 and below — but only where you supply the language model or the voice and pay that vendor separately, so they are not like-for-like. For narration, cheapest is rarely the right question — voice quality and commercial rights usually decide it.
Is this ranking based on hands-on testing?
No. Every price and rate is quoted from an official vendor pricing page checked in September 2026, and every derived figure is CartSignal arithmetic over those published numbers. No audio was generated, and no voice quality, naturalness or pronunciation accuracy is assessed here.
Sources and related pages
Rates come from ElevenLabs' pricing page and API pricing page, Deepgram's pricing page, Cartesia's pricing page and Inworld's pricing page, all checked 13 September 2026, together with the Murf and Speechify rate cards recorded on their CartSignal profiles (verified 30 August and 10 September 2026). Murf's pricing page rendered no readable figures on fetch, so its Studio and API rates are carried from the Murf profile rather than re-verified today.
Related: ElevenLabs alternatives, ranked per character · Voice compare · ElevenLabs review · Cartesia review · Murf review · Speechify review · AI voice category · Best AI video tools · EU AI Act labelling · AI-readable feed.