ElevenLabs
AI voice generation, text-to-speech, dubbing, voice cloning, and voice agents for creators and businesses.
Focused CartSignal directory category for ai voice used by creators and small businesses. Built as a focused vertical, not broad AI-tool sprawl.
| Tool | Positioning | Best fit | Review basis |
|---|---|---|---|
| ElevenLabs | AI voice generation, text-to-speech, dubbing, voice cloning, and voice agents for creators and businesses. | Creator voiceovers, Podcast narration | Public-data review |
| Cartesia | Real-time voice AI built on state space models: Sonic text-to-speech, Ink transcription and managed voice agents, all billed from one credit pool. The cheapest per-minute TTS listed here at about $0.038 on the $5 Pro plan. | Low-latency voice apps, Agent voice interfaces | Public-data review |
| PlayHT | Discontinued. Meta acquired the team in July 2025 and the service was withdrawn; play.ht no longer resolves. See what to use instead. | Not available — do not buy | Discontinued listing |
| Resemble AI | No longer sells voice AI. Resemble announced on 13 August 2026 that it is "not selling voice AI to new customers" and moved the whole company to deepfake detection, watermarking and identity matching. Its voice models stay free under the MIT licence as Chatterbox. See what changed. | Deepfake detection, Watermarking and provenance — not hosted voice | Public-data review |
| Murf AI | Browser voiceover studio billed in Voice Generation Time, plus a separately priced API (Falcon 2, Gen 2) and Murf Dub video localisation in 40+ languages. | Training and e-learning narration, Team voiceover workflows | Public-data review |
| Speechify | Three separately billed products: a $29/month reading app that generates nothing publishable, Speechify Studio at $100/year on credits (1/sec voiceover, 3 dubbing, 30 avatars), and a developer API about 10x cheaper per spoken minute than Studio — the lowest published TTS rate listed here. | High-volume programmatic narration, Budget dubbing and avatar video | Public-data review |
| Deepgram | Developer-first speech platform sold by usage rather than subscription: Nova-3 and Flux transcription from $0.258 an hour, Aura and Flux TTS on a clean $15/$30/$45 per-million-character ladder, and a bundled voice agent that costs about 6x the sum of its own component rates. The 15 September 2026 increase took the standard agent minute to $0.075 but left every bring-your-own-TTS row untouched. | Production voice agents, High-volume transcription | Public-data review |
| OpenAI audio API | Voice sold only as a developer API — no studio, no subscription, no free tier — across four meters: text to speech from $15.00 per million characters, file transcription at $0.27 an hour, streaming transcription at $1.02, and speech-to-speech conversation billed per audio token. The realtime meter has no flat per-minute rate, because the whole conversation is re-billed as input on every turn. | Programmatic narration, Bulk file transcription | Public-data review |
| Amazon Polly | Text to speech sold purely by usage inside AWS, with no plan, credits or minimum. Four engines at $4.00, $16.00, $30.00 and $100.00 per million characters — the cheapest published rate in this directory and the dearest one too. Price and coverage run in opposite directions: $4.00 Standard reaches 29 language variants everywhere, $100.00 Long-form is six voices in two languages in one region. Word-level timing is billed as a second pass, so lip sync doubles the card. | High-volume narration, IVR and accessibility read-along | Public-data review |
| AssemblyAI | Speech-to-text, speech understanding and a bundled voice agent, sold by usage with no plan. Transcription runs $0.15 to $0.45 an hour and the agent bundle is $4.50 — 30x its own cheapest streaming rate. Two different meters: pre-recorded audio is billed on file duration to the exact second, while streaming is billed on how long the socket stays open, idle time included. The $0.21 flagship covers 18 languages; the $0.15 model beneath it covers 99. | High-volume transcription, Live captions and voice agents | Public-data review |
AI voice generation, text-to-speech, dubbing, voice cloning, and voice agents for creators and businesses.
Real-time voice AI built on state space models: Sonic text-to-speech, Ink transcription and managed voice agents, all billed from one credit pool. The cheapest per-minute TTS listed here at about $0.038 on the $5 Pro plan.
Discontinued. Meta acquired the PlayAI team in July 2025, the Voice API and TTS service were deprecated within weeks, and the play.ht domain no longer resolves. Kept as a correction because older guides still recommend it.
No longer sells voice AI to new customers. Resemble moved the company to deepfake detection and watermarking in August 2026; its pricing page carries no text-to-speech rate. The voice models remain free and self-hostable under the MIT licence as Chatterbox.
Browser voiceover studio billed in Voice Generation Time, plus a separately priced API (Falcon 2, Gen 2) and Murf Dub video localisation in 40+ languages.
Three products on three unrelated meters. The reading app most buyers find first costs $29/month and produces nothing publishable; Speechify Studio ($100/year) does the commercial work on credits that burn 30x faster for avatars than voiceover; the developer API is the cheapest published TTS rate on this site.
Sold by usage, not by plan, starting from a $200 credit with no expiry. Transcription from $0.258 an hour, text to speech on a 1:2:3 ladder at $15, $30 and $45 per million characters, and a bundled voice agent minute that costs roughly 6x Deepgram's own speech rates combined. Agent rates rose a third on 15 September 2026 — but only on the rows using Deepgram's own voice.
One rate card, no plans, no free tier. Speech from $0.0144 a minute and file transcription at $0.27 an hour put it at the cheap end — but streaming transcription is the dearest listed here at $1.02, its realtime agent meter compounds with call length, and the only model that returns subtitle timestamps is scheduled for removal in February 2027.
Text to speech sold purely by usage inside AWS, with no plan, credits or minimum. Four engines at $4.00, $16.00, $30.00 and $100.00 per million characters — the cheapest published rate in this directory and the dearest one too. Price and coverage run in opposite directions: $4.00 Standard reaches 29 language variants everywhere, $100.00 Long-form is six voices in two languages in one region. Word-level timing is billed as a second pass, so lip sync doubles the card.
Transcription and voice agents sold by usage, from $0.15 an hour, with $50 of free credit and the most complete add-on card in this directory. The catch is not the rate: streaming is billed on connection time rather than audio, so a session carrying speech half the time costs 2x its published rate and one nobody closes bills for three hours.