ElevenLabs Explained: Text-to-Speech, Voice Cloning, and Dubbing
What the platform does, what it costs, and where it beats or loses to OpenAI TTS, PlayHT, and Cartesia.
ElevenLabs is an AI audio platform that converts text into realistic speech, clones a person's voice from a short recording, and dubs video into dozens of languages while keeping the original speaker's tone.
The company started as a text-to-speech lab and has grown into a full audio stack: TTS, speech-to-text, voice cloning, dubbing, sound effects, and conversational voice agents.
This guide covers the model lineup, the Voice Library, API versus dashboard workflows, current pricing, and where ElevenLabs wins or loses against OpenAI TTS, PlayHT, and Cartesia.
What Does ElevenLabs Actually Do
ElevenLabs turns written text into natural-sounding speech, and it turns a short voice sample into a reusable, clonable voice.
The core product is text-to-speech (TTS): paste in text, pick a voice, and get back an audio file. Under that sits three more capabilities that separate it from a basic TTS API.
Voice cloning builds a synthetic version of a real voice from as little as one minute of audio. Dubbing Studio translates a video's audio into another language while syncing lip movement and preserving the speaker's vocal character. Speech-to-Text (Scribe) transcribes audio back into text for captioning or editing workflows.
Most teams start in the web dashboard, then move production workloads to the API once a voice and script format are locked in.
- Text-to-Speech: text in, natural speech out, 70+ languages depending on model
- Voice cloning: Instant Voice Cloning (~1 minute of audio) or Professional Voice Cloning (longer samples, higher fidelity)
- Dubbing Studio: translates video/audio and syncs the dub to lip movement
- Speech-to-Text (Scribe): transcription for captions and content repurposing
- Conversational AI: low-latency voice agents built on the same voice stack
Not sure whether ElevenLabs, OpenAI TTS, or Cartesia fits your voice AI use case? Get a straight answer.
Book a ConsultationWhich ElevenLabs Model Should You Use
ElevenLabs ships different models for different jobs: expressiveness, speed, or raw language coverage.
Eleven v3 is the most expressive model. It supports inline audio tags like [whispers] and [excited], multi-speaker dialogue in a single generation, and wide language coverage, which makes it the pick for audiobooks, character voices, and narrative content.
Eleven Multilingual v2 is the steady, production-grade default for straightforward narration and voiceover work.
Flash-tier models trade some expressiveness for very low latency, which is why they show up in real-time voice agents and live conversational apps rather than pre-rendered video narration.
Model names and specific latency numbers change as ElevenLabs ships updates, so confirm the current model list and per-model limits on the vendor's models page before locking in a production pipeline.
- Eleven v3: most expressive, audio tags, multi-speaker dialogue, widest language coverage
- Eleven Multilingual v2: reliable default for standard narration and voiceover
- Flash-tier models: built for low-latency, real-time voice agent use cases
- Turbo-tier models: a middle ground between expressiveness and speed
How the Voice Library and Voice Cloning Work
The Voice Library is a searchable catalog of thousands of pre-made voices, and cloning lets you add your own on top of it.
Instant Voice Cloning (IVC) needs about a minute of clean audio and produces a usable clone in minutes. It is available starting on the Starter plan.
Professional Voice Cloning (PVC) trains on a longer sample set for a closer, higher-fidelity match, and it unlocks on the Creator plan and above.
Voice Design generates an entirely new synthetic voice from a text description instead of a recording, useful when no real sample exists yet.
- Voice Library: thousands of community and stock voices, filterable by accent, age, and use case
- Instant Voice Cloning: fast, ~1 minute of source audio, Starter plan and up
- Professional Voice Cloning: higher fidelity, longer training sample, Creator plan and up
- Voice Design: generate a new voice from a text description, no source recording needed
How Dubbing Studio and Speech-to-Text Work
Dubbing Studio translates a video's spoken audio into another language and syncs the new audio to the speaker's lip movement.
You upload a video or audio file, pick a target language, and Dubbing Studio produces a translated track that keeps the original speaker's vocal tone rather than swapping in a generic voice.
The API version of dubbing accepts much larger files than the dashboard UI, which matters for long-form video or podcast catalogs.
Scribe, the speech-to-text model, runs the reverse direction: audio in, transcript out, used for captions, editing transcripts, or feeding a script back into TTS.
- Dubbing Studio: video/audio translation with lip-sync and voice-tone preservation
- Coverage: dozens of languages; confirm the current list on the vendor's dubbing page
- API dubbing endpoint: accepts far larger files than the dashboard UI allows
- Scribe (Speech-to-Text): transcription for captions and content repurposing
API vs Dashboard: Which One Should You Use
Use the dashboard to design and test voices, and use the API to generate audio at scale in a product or pipeline.
The dashboard is the fastest way to browse the Voice Library, run Voice Design, and preview a script before committing credits to it.
The API is built for automation: generating narration for hundreds of articles, powering a voice agent, or batch-dubbing a video catalog.
Teams that skip the dashboard and go straight to the API often waste credits on trial-and-error voice selection they could have done for free in the preview tool first.
- Dashboard: voice discovery, script testing, Dubbing Studio uploads, account management
- API: production TTS generation, programmatic dubbing, Conversational AI agents
- Both draw from the same credit pool, so plan usage across both surfaces together
How Much Does ElevenLabs Cost
ElevenLabs runs on a credit-based subscription, from a free tier through custom Enterprise pricing.
The Free plan includes 10,000 credits a month, enough for a short test project but no commercial license.
Starter unlocks commercial rights, Instant Voice Cloning, and Dubbing Studio for a low monthly price. Creator adds Professional Voice Cloning. Pro and above raise the credit ceiling and unlock higher-quality audio output for production teams.
Business-tier plans add more seats, more Professional Voice Clones, and lower-latency TTS for real-time products. Enterprise adds custom terms, HIPAA BAAs, and custom SSO.
Prices and credit allotments change as ElevenLabs adjusts plans, so treat any number here as a starting point and verify current pricing on ElevenLabs' own pricing page before budgeting.
- Free: 10,000 credits/month, no commercial license
- Starter: commercial license, Instant Voice Cloning, Dubbing Studio unlock
- Creator: adds Professional Voice Cloning, higher credit ceiling
- Pro: higher-quality audio output, larger credit pool for production use
- Scale / Business: more seats, more Professional Voice Clones, lower-latency TTS
- Enterprise: custom pricing, HIPAA BAAs, custom SSO
How ElevenLabs Compares to OpenAI TTS, PlayHT, and Cartesia
ElevenLabs generally leads on voice cloning quality and language breadth, while Cartesia leads on raw latency and OpenAI's TTS API wins on simplicity for teams already inside the OpenAI ecosystem.
OpenAI's TTS models are easy to adopt if you already call the OpenAI API for text, but the voice library and cloning controls are far more limited than ElevenLabs.
PlayHT competes closely on cloning quality and offers competitive per-character pricing at high volume, though most independent listening comparisons still rate ElevenLabs' emotional range higher, especially on the v3 model.
Cartesia focuses on ultra-low-latency streaming for real-time voice agents. If your product's whole value is sub-100ms conversational response time, Cartesia and ElevenLabs' flash-tier models are the two to benchmark head-to-head with your own audio, not a published leaderboard.
Published TTS benchmark scores age fast and vary by listening panel, so run a blind A/B test on your own scripts before picking a vendor based on a marketing chart.
- ElevenLabs: broadest voice library, strongest expressive/emotional range (v3), widest language coverage
- OpenAI TTS: simplest to integrate if already on the OpenAI API, smaller voice/cloning toolset
- PlayHT: competitive cloning quality, strong at high-volume pricing
- Cartesia: built for the lowest streaming latency in real-time voice agents
What Are the Consent and Licensing Rules for Voice Cloning
You need explicit permission from a person before cloning their voice, and ElevenLabs' terms require you to hold the rights to any voice you upload.
The Free and lower tiers restrict commercial use of generated audio; a paid plan's commercial license only covers voices you have the right to use, not any voice you can technically clone.
Cloning a public figure's voice without consent, even for satire or a demo, carries real legal exposure in most jurisdictions and can trigger platform account action on top of that.
The safest pattern for business use is a licensed stock voice from the Voice Library or a cloned voice with a signed release from the speaker.
- Get written consent before cloning anyone's real voice
- Commercial license only covers voices you actually have rights to use
- Public-figure voice cloning without consent is a legal and platform risk, not just an ethics question
- Stock Voice Library voices carry cleaner licensing than a cloned voice with no release
Where ElevenLabs Voice Projects Go Wrong
Most ElevenLabs failures trace back to script formatting, model mismatch, or credit budgeting, not the underlying voice quality.
A script with abbreviations, unusual proper nouns, or numbers formatted inconsistently will mispronounce more often than a clean, expanded script.
Picking Eleven v3 for a real-time agent, or a flash-tier model for an emotional audiobook chapter, produces the wrong trade-off in both directions: v3 adds latency an agent cannot afford, and flash-tier models sound flatter than v3 on narrative content.
When we run our GA4 top-pages routine across the client sites in our own portfolio, the pages that hold up longest are the ones written in short, plainly formatted sentences before any AI layer touches them - the same discipline that keeps a TTS script from mispronouncing itself. Clean input text is a cheap fix that prevents a re-render and re-spent credits.
Credit budgets run out mid-project when teams test extensively in the API instead of the free dashboard preview, since both draw from the same monthly pool.
- Mispronunciations usually trace to messy source text, not the model
- Model mismatch: v3 for real-time agents, flash-tier for emotional narration, both are wrong-way picks
- Credit burn: testing in the API instead of the free dashboard preview wastes budget fast
- Dubbing sync issues usually come from source video with overlapping speech or background music
Frequently Asked Questions
- Yes, ElevenLabs offers a Free plan with 10,000 credits a month, but it does not include a commercial license, so output cannot be used in paid or public commercial projects. Paid plans start with the Starter tier, which adds commercial rights, Instant Voice Cloning, and Dubbing Studio.
- ElevenLabs voice cloning works by training a model on a sample of a person's voice, then generating new speech in that voice from text input. Instant Voice Cloning needs about a minute of clean audio, while Professional Voice Cloning trains on a longer sample for higher fidelity and requires the Creator plan or above.
- ElevenLabs generally offers a larger voice library, stronger voice cloning, and more expressive output than OpenAI's TTS models, especially with the Eleven v3 model. OpenAI TTS is simpler to adopt if a team already uses the OpenAI API for text generation, but its cloning and voice-customization options are more limited.
- Only with that person's explicit consent. ElevenLabs' terms require you to hold rights to any voice you clone, and cloning a real person's voice without permission carries legal exposure regardless of the platform's technical capability. Stick to voices you own the rights to or voices with a signed release.
- ElevenLabs' Dubbing Studio supports dozens of languages, and the exact list changes as the company adds coverage. Check the current language list on ElevenLabs' own dubbing page before committing to a project that depends on a specific language pair.
- Flash-tier models prioritize very low latency for real-time use cases like voice agents, while Eleven v3 prioritizes expressiveness, with features like inline emotion tags and multi-speaker dialogue. Pick Flash for anything a user talks to live, and v3 for pre-rendered narrative content like audiobooks or ads.
- ElevenLabs pricing runs from a free tier through custom Enterprise pricing, with paid individual plans in between covering commercial use, voice cloning tiers, and rising credit allotments. See our [ElevenLabs pricing guide](/guides/elevenlabs-pricing) for the current tier breakdown, and verify exact numbers on ElevenLabs' own pricing page.
Not Sure Which Voice AI Stack Fits Your Product
ElevenLabs, OpenAI TTS, PlayHT, and Cartesia solve different problems well. We help teams pick the right voice AI stack for their actual use case, wire it into a production pipeline, and avoid the credit-budget and licensing mistakes that show up after launch.
Book a Consultation