Skip to content
GuidesSeptember 30, 2026· 8 min read

ElevenLabs API Pricing: What a Voiceover Really Costs (2026)

ElevenLabs API pricing per 1,000 characters by model, what a 30s or 60s voiceover costs, and how to batch-generate voiceovers from code without 429 errors.

elevenlabsvoiceoverttspricing
Guides

ElevenLabs API Pricing: What a Voiceover Really Costs (2026)

A 30-second voiceover is about 600 characters of script. On ElevenLabs' cheapest standard API rate at the time of writing, that is roughly two and a half cents of audio. The number that actually decides your bill is how many takes, languages and variants you render, and which model you picked without thinking about it.

This guide works through the per-character rates model by model, converts them into the durations you actually ship (30 and 60 seconds), prices realistic batches, and then covers the part most pricing pages skip: generating hundreds of voiceovers from code without tripping the concurrency limit. All figures were checked on 30 September 2026. ElevenLabs changes its rates, sometimes with time-limited promotions, so treat the live pricing page as the final word.

ElevenLabs API pricing per model, as of September 2026

ElevenLabs bills text-to-speech by character. Its API pricing page lists these per-1,000-character rates at the time of writing:

  • Eleven v4: $0.022 per 1K characters, a 72% promotional discount from the $0.08 list price, running until 12 October.
  • Eleven v4 Turbo: $0.011 per 1K characters, also 72% off, from a $0.04 list price, same end date.
  • Eleven v3: $0.08 per 1K characters.
  • Eleven v3 Conversational: $0.04 per 1K characters.
  • Multilingual v2: $0.08 per 1K characters.
  • Flash and Turbo: $0.04 per 1K characters.

Two things to note before you build a budget on those numbers. First, the v4 prices are promotional. Any cost model that runs past mid-October should use the $0.08 and $0.04 list prices, or it will understate spend by close to four times. Second, older third-party write-ups quote different figures. A June 2026 Puter tutorial lists $0.05 per 1K for Flash and Turbo and $0.10 per 1K for Multilingual v2 and v3. Rates have moved within a single quarter, which is the main argument for pulling prices from the source at run time instead of hard-coding them.

How credits map to characters

The consumer pricing page meters text-to-speech at 1 credit per character. Flash models are cheaper in credit terms: Smallest.ai's breakdown puts them at 0.5 to 1 credit per character depending on plan. That second figure comes from a third party, so confirm it against your own usage dashboard before you rely on it.

What a 30-second and a 60-second voiceover cost

Rates are quoted per 1,000 characters, while scripts get planned in seconds, so you need a conversion. The usual one is about 1,200 characters per minute of speech, which Smallest.ai derives from a two-minute narration landing around 2,400 characters. That gives roughly 600 characters for a 30-second read and 1,200 for 60 seconds. A 500-word script lands near 3,000 characters by the same source.

Apply the September 2026 API rates:

  • Flash / Turbo ($0.04 per 1K): 30s = 0.6 x $0.04 = $0.024. 60s = 1.2 x $0.04 = $0.048.
  • Multilingual v2 or v3 ($0.08 per 1K): 30s = $0.048. 60s = $0.096.
  • v4 at promo ($0.022 per 1K): 30s = about $0.013. 60s = about $0.026.
  • v4 Turbo at promo ($0.011 per 1K): 30s = about $0.007. 60s = about $0.013.

A cost formula you can paste into a spreadsheet

At about 1,200 characters per minute, speech runs near 20 characters per second. That gives a one-line estimate:

  • Characters = seconds x 20
  • Cost per voiceover = characters x rate per 1K / 1,000
  • Cost per batch = cost per voiceover x variants x languages x takes

The last line is the one people forget. Three takes per script across two languages multiplies the voice line by six before anyone listens to a single file.

Pacing moves these numbers. A calm product explainer reads slower than a fast UGC-style hook, so the same 30 seconds can hold noticeably more or fewer characters. If you are still deciding on tone and speed, the trade-offs are covered in picking the right voice and pacing for ad voiceovers. For budgeting, count characters in the actual script and treat the per-minute rule as a sanity check.

Subscription credits versus the per-character API rate

ElevenLabs also sells monthly plans that bundle credits. The pricing page lists, at the time of writing:

  1. Free: 10,000 credits, no commercial use.
  2. Starter: 30,000 credits for $6.
  3. Creator: 121,000 credits for $22 ($11 for the first month).
  4. Pro: 600,000 credits for $99.
  5. Scale: 1,800,000 credits for $299.
  6. Business: 6,000,000 credits for $990.

Divide price by credits and you get the implied cost per 1,000 credits: $0.20 on Starter, about $0.18 on Creator (about $0.09 in the discounted first month), and about $0.165 on Pro, Scale and Business. At 1 credit per character, those implied rates sit above the per-character API list prices quoted earlier. Puter's tutorial describes the model as the same per-unit rate across tiers, with subscriptions bundling usage. The two official pages do not reconcile neatly from the outside, so before committing to a plan, run a week of real traffic and compare the invoice to the arithmetic.

What plans do buy you, beyond credits, is headroom. The plan tier sets your concurrency ceiling (next sections) and the commercial licence. ElevenLabs states that commercial use starts at Starter and the text-to-speech docs repeat that commercial use requires a paid plan. Anything that runs in paid media needs at least Starter.

If you are weighing bundled credits against paying per call in general, the reasoning carries over from video and images, and it is laid out in credits versus pay-per-call pricing.

Batch budgets: the arithmetic for real workloads

A single voiceover costs too little to matter, so model choice only shows up once you render in batches. Here are four common workloads, priced at the September 2026 API list rates for Flash ($0.04 per 1K) and Multilingual v2 ($0.08 per 1K), using 600 characters per 30 seconds.

20 hook variants for one ad

Hooks are short. Say each hook plus body is a 30-second read, so 20 variants x 600 characters = 12,000 characters. Flash: $0.48. Multilingual v2: $0.96. If you are writing those variants, these hook templates give you 18 openers to rotate through.

A week of daily 60-second shorts

7 x 1,200 = 8,400 characters. Flash: about $0.34. Multilingual v2: about $0.67.

50 product explainers at 60 seconds

50 x 1,200 = 60,000 characters. Flash: $2.40. Multilingual v2: $4.80.

One 60-second ad localised into six languages

6 x 1,200 = 7,200 characters. Multilingual v2: about $0.58. Translated scripts rarely keep the exact English length, so budget per language from the translated text. The production side of this is covered in running video ads in several languages.

At volume the gap between models is real but small in absolute terms. Puter's example of 3 million characters came to $150 on Flash against $300 on Multilingual, at its older rates. For most ad and content teams the voice line is the cheapest part of the stack. The expensive mistake is re-rendering whole scripts to fix one sentence, which the continuity parameters below help you avoid.

Choosing a model: price, languages, limits and latency

Price is one axis of four. The ElevenLabs model docs list the others:

  • Languages: v4 covers 90+, v3 covers 70+, Flash v2.5 covers 32, Multilingual v2 covers 29, and Flash v2 is English only.
  • Characters per request: Flash v2.5 accepts 40,000, Flash v2 30,000, v4 and Multilingual v2 10,000 each, and v3 5,000.
  • Latency: Flash v2.5 around 75ms, v4 Turbo around 100ms, v3 Conversational around 280ms.

A decision rule you can reuse

  1. English only, high volume, latency matters (live agents, previews, internal drafts): Flash. Lowest list price, largest request size, fastest.
  2. Final ad voiceover in one of the major languages: Multilingual v2 or v3 at list price, or v4 while the promotion runs. Render a few takes and keep the best.
  3. Language outside Multilingual v2's 29: check v4 (90+) or v3 (70+) coverage first.
  4. Long-form narration in one call: Flash v2.5's 40,000-character limit fits a script that would need splitting on v3's 5,000.
  5. Budget beyond 12 October: model v4 at $0.08 and v4 Turbo at $0.04, whatever the page shows today.

Quality judgements between models are subjective and change with each release. The practical test is cheap: at a few cents per 60 seconds, render the same script on two models and listen.

Generating voiceovers in batch from code

The endpoint is simple. The API reference documents it as follows:

  • Request: POST https://api.elevenlabs.io/v1/text-to-speech/{voice_id}, with your key in the xi-api-key header.
  • Body: text is required. model_id defaults to eleven_multilingual_v2, and output_format defaults to mp3_44100_128.
  • Consistency: seed takes an integer from 0 to 4,294,967,295, and previous_text and next_text give the model context around a fragment.
  • Normalisation: apply_text_normalization accepts auto, on or off.
  • Response: binary audio on 200, and 422 on a validation error.

Note the default model. If your script omits model_id, every call runs on Multilingual v2 at $0.08 per 1K, twice the Flash rate. Set it explicitly.

Output format is a cost-free decision you should still make

The text-to-speech docs list MP3 at 22.05 to 44.1kHz and 32 to 192kbps, PCM at 16 to 44.1kHz, mu-law and A-law at 8kHz for telephony, and Opus at 48kHz. For audio that goes into a video editor, PCM avoids a second lossy encode. For web previews, the MP3 default is fine.

A batch job, step by step

  1. Prepare a manifest. One row per output: id, voice_id, model_id, text, seed, output path. A row looks like hook-07-en, your voice id, the Flash model id, "Your first draft took three hours...", 42, out/hook-07-en.mp3. Count characters per row and sum them. That sum times the per-1K rate is your cost ceiling before you send anything.
  2. Split long text on sentence boundaries below the model's per-request limit, and pass the neighbouring chunks as previous_text and next_text so the joins sound continuous, as the docs recommend.
  3. Fix the seed per voice. Re-rendering one line with the same seed and settings keeps it closer to the rest of the take, so you fix one sentence instead of paying for the whole script again.
  4. Run a bounded worker pool sized to your plan's concurrency (next section), so requests never exceed the ceiling.
  5. Write each file and a log line with id, model, character count and computed cost. That log is how you reconcile against the invoice.
  6. Retry only what failed, keyed by id, so a crash halfway through never re-bills finished rows.

Concurrency is the limit that bites

ElevenLabs meters concurrent requests, not requests per minute, according to its rate limiting write-up. The ceilings by plan, from the 429 error help page:

  • Free: 2
  • Starter: 3
  • Creator: 5
  • Pro: 10
  • Scale and Business: 15

Flash models get higher limits, between 4 and 30 depending on plan. Requests above the ceiling queue by plan priority, which ElevenLabs says adds roughly 50ms, and past that you get a 429.

Handling 429s correctly

The help page distinguishes two 429 causes, and they need different handling:

  • too_many_concurrent_requests: you exceeded your plan's concurrency. Retrying immediately makes it worse. Shrink the pool.
  • system_busy: load on ElevenLabs' side. It usually succeeds on retry.

The recommended pattern is a bounded concurrency pool, a token bucket, and exponential backoff with jitter. Responses carry current-concurrent-requests and maximum-concurrent-requests headers, so your pool can read its real ceiling instead of trusting a config value.

Batch runner checklist

  • Pool size set from maximum-concurrent-requests, minus one if anything else shares the key.
  • model_id set explicitly on every call.
  • Backoff with jitter on system_busy; pool shrink on too_many_concurrent_requests.
  • No retry on 422: fix the payload instead.
  • Character count and cost logged per file.
  • Idempotent ids so a rerun skips completed rows.
  • Promo-period prices flagged in the cost log, so next month's numbers do not surprise anyone.

At a pool of 5 on Creator, 50 explainers finish in roughly ten rounds of requests. The wall-clock time is dominated by generation, and the job is small enough to run from a laptop.

Voiceover inside a larger pipeline

A voiceover is rarely the deliverable. It sits next to a video clip, a product shot and captions, and each of those comes from a different model with its own key, bill and rate limit. Teams that produce at volume usually end up driving the whole chain from a script or an agent: a manifest in, files and their costs out. A working version of that setup, with an agent orchestrating the steps, is described in automating video production with AI agents, and the Claude Code side specifically in generating video from Claude Code through an MCP server.

The cost logic is the same at every step. Count units (characters, seconds, images), multiply by the current per-unit price, log it next to the file, and reconcile. Voice is the easiest line to get right, because characters are known before you send the request.

FAQ

How much does the ElevenLabs API cost per character?

At the time of writing, the API pricing page lists $0.04 per 1,000 characters for Flash and Turbo, $0.08 for Multilingual v2 and v3, and promotional rates of $0.022 for v4 and $0.011 for v4 Turbo until 12 October. That is $0.00004 to $0.00008 per character at list price.

How much does a 1-minute AI voiceover cost with ElevenLabs?

About 1,200 characters, so roughly $0.048 on Flash and $0.096 on Multilingual v2 at September 2026 API rates. Your script's actual character count is the precise input.

Is there a free ElevenLabs API tier?

The Free plan includes 10,000 credits per month and 2 concurrent requests, but it does not permit commercial use. Voiceovers for paid campaigns or client work need Starter or above.

Why am I getting 429 errors from the ElevenLabs API?

You are either over your plan's concurrent request limit (too_many_concurrent_requests) or ElevenLabs is under load (system_busy). Reduce your worker pool for the first, retry with backoff for the second, as the help page explains.

What is the maximum text length per ElevenLabs API request?

It depends on the model: 40,000 characters on Flash v2.5, 30,000 on Flash v2, 10,000 on v4 and Multilingual v2, and 5,000 on v3, per the model docs.

Sources

  1. ElevenLabs: API Pricing
  2. ElevenLabs: Pricing
  3. ElevenLabs: Models documentation
  4. ElevenLabs: Text to Speech capability overview
  5. ElevenLabs: Text to Speech convert API reference
  6. ElevenLabs: API Error Code 429
  7. ElevenLabs: AI rate limiting for voice
  8. Puter: ElevenLabs API Pricing, Full Breakdown of Costs (Jun 2026)
  9. Smallest.ai: ElevenLabs Pricing Explained

If you would rather not hold a separate ElevenLabs subscription next to your video and image keys, Aitachyon runs ElevenLabs voiceover from the same prepaid balance as Seedance, Kling, Veo and the image models, at $0.043 to $0.086 per 450 characters at the time of writing (see the ElevenLabs model page). That per-character rate is above ElevenLabs' direct API list price. In exchange you get one key for the web studio, the HTTP API and the hosted MCP server, every job itemised with the model it ran on and its cost, automatic refunds on failed renders, a per-key spend alert, and a balance that never expires. Each voiceover gets a stable vo_ ref your code can fetch with GET /api/generations/{ref}. The API and MCP quick start has the one-line Claude Code setup.

Related articles

Free tools to try

Stop describing your brand. Paste your URL.

Aitachyon reads your whole brand from your website, then creates videos, images, carousels, posts and banners, on-brand, every format, every feed.