Workspaces are coming: shared projects, assets and credits for your whole team

Follow along
Audio

Paste the script. Choose how it is read.

Nine voice models and forty-six named voices behind one field, with real controls over stability, emotion, pitch and pace.

Real output

Scripts and the reads they produced

Speech has no thumbnail. Each entry is the text that was pasted and the waveform of what came back.

  • Prompt

    “Every generation you run lands in your library, tagged and searchable, from the moment it finishes.”

    0:08

    Product narration

    ElevenLabs v3 on Aria at stability 0.6.

  • Prompt

    “I didn't think it would work. It worked.”

    0:04

    Directed line

    Minimax Speech 2.8 HD on Wise Woman, surprised.

  • Prompt

    “Take your time with this one. There is no rush, and nothing here is going anywhere.”

    0:07

    Lowered read

    Minimax on Deep Voice Man with pitch at −4 and speed at 0.9.

  • Prompt

    “Cada generación aparece en tu biblioteca, etiquetada y lista para buscar.”

    0:06

    Spanish voice-over

    Qwen3 TTS 1.7B on Vivian, with the language detected automatically.

How it works

A script, a voice, and the delivery

01

Paste your script

Up to 5000 characters in one run, enforced as you type. These models read punctuation as direction: a comma is a breath and a full stop is a beat — write the pauses in rather than asking for them.

02

Pick a model, then a voice

Nine models across four families. ElevenLabs brings twenty named voices, Minimax seventeen and Qwen nine; Cartesia offers exactly one voice and no picker, because there is nothing to choose between.

03

Tune the read and generate

Speed on every family, plus stability, similarity and style on ElevenLabs or pitch and emotion on Minimax. The audio lands in your workspace library, and failed runs are refunded.

Inputs and outputs

What Text to Speech takes, and what it returns

What Text to Speech needs from you

Script
Text, up to 5000 characters per run, enforced as you type
Voice
One of 46 named voices — 20 on ElevenLabs, 17 on Minimax, 9 on Qwen, and one fixed voice on Cartesia
Price
Charged per 1000-character block, with a one-block floor — a short line costs a full block
Languages
Qwen3 TTS covers eleven with automatic detection; ElevenLabs Multilingual v2 holds one voice across thirty

Export specifications

Per run
Up to 5000 characters of script
Models
Nine, across ElevenLabs, Minimax, Qwen and Cartesia
Voices
46 named voices, plus Cartesia's single fixed voice
Price
Per 1000-character block, at the rate of the model you chose, with a one-block floor
Failed runs
Refunded
Delivery
Rendered audio file in your workspace library
What you can control

Not just which voice — how it reads the line

The controls differ by model, because they are the model's own parameters rather than a common wrapper. Speed is the one every family offers.

  • Script

    5000 characters per run

    Enforced as you type, and priced per 1000-character block with a one-block floor. Longer scripts split into runs — break at a paragraph so the phrasing does not fall apart at the seam.

  • ElevenLabs voices

    20 named voices, Aria the default

    Aria, Roger, Sarah, Laura, Charlie, George, Callum, River, Liam, Charlotte, Alice, Matilda, Will, Jessica, Eric, Chris, Brian, Daniel, Lily and Bill.

  • Minimax voices

    17 named voices, Wise Woman the default

    Wise Woman, Friendly Person, Inspirational Girl, Deep Voice Man, Calm Woman, Casual Guy, Lively Girl, Patient Man, Young Knight, Determined Man, Lovely Girl, Decent Boy, Imposing Manner, Elegant Man, Abbess, Sweet Girl 2 and Exuberant Girl.

  • Qwen voices

    9 named voices, Vivian the default

    Vivian, Serena, Uncle Fu, Dylan, Eric, Ryan, Aiden, Ono Anna and Sohee. Cartesia is the exception: it offers exactly one voice and no picker.

  • Speed

    0.7–1.2 on ElevenLabs, Qwen and Cartesia; 0.5–2 on Minimax

    The one control every family offers. The narrow band is narrow on purpose — pushed further, the read stops sounding like a person. Minimax gives the wider range, for anything from a slow read to a compressed disclaimer.

  • Stability

    0–1, default 0.5

    ElevenLabs only. Low is expressive and varies between takes; high is consistent and flatter.

  • Similarity

    0–1, default 0.75

    ElevenLabs only. How closely the render holds to the reference voice.

  • Style

    0–1, default 0

    ElevenLabs only. Pushes the delivery away from neutral. Zero is the default for a reason.

  • Pitch

    −12 to +12, default 0

    Minimax only. Semitone shift on the voice, without changing the pace.

  • Emotion

    Neutral, happy, sad, angry, fearful, disgusted, surprised

    Minimax only. Seven named settings, chosen rather than coaxed out of the punctuation.

  • Language

    Eleven on Qwen3 TTS, with automatic detection

    Auto, English, Chinese, Spanish, French, German, Italian, Japanese, Korean, Portuguese and Russian.

Engine choice

Pick the one that matches your source.

  • ElevenLabs v3

    Expressive

    The most expressive. Reads emotion out of the writing.

    Narration that has to carry feeling

  • ElevenLabs Multilingual v2

    Multilingual

    Thirty languages, one voice across all of them.

    Scripts that switch language

  • ElevenLabs Turbo v2.5

    Cheaper

    Half the price and most of the quality. Good for drafts.

    Drafts and volume

  • Minimax Speech 2.8 Turbo

    Controls

    Emotion, pitch and volume as real controls rather than hints.

    Directed performance, quickly

  • Minimax Speech 2.8 HD

    Fidelity

    The same voices, rendered at higher fidelity.

    A finished directed read

  • Qwen3 TTS 1.7B

    Languages

    Eleven languages with automatic detection.

    Non-English scripts

  • Qwen3 TTS 0.6B

    Fast

    The smaller, faster one. Same voices.

    Iterating on a non-English script

  • Cartesia Sonic 3.5

    Instant

    Finishes while you wait — nothing to come back for.

    A read you need now

  • Minimax Voice Clone

    Cloning

    Reads your script in a voice from a recording you supply.

    A voice of your own

Who it's for

Built for anyone who needs a script read aloud

  1. Video creators

    Narrate a cut without booking a booth, and re-render the line when the script changes.

  2. Course and training teams

    Keep one named voice across dozens of modules, including the ones written months apart.

  3. Product teams

    Voice a demo or an in-app walkthrough, in more than one language, from the same script.

One product. Pick your volume.

Built for Growth at Every Stage

Every plan includes every model and every feature. Plans only change how many credits you get and how many generations run at once.

Pro

Occasional projects

2,900/month

Get Started

Credits per month

60,000

≈1,000 images or ≈12 videos


At once

6 parallel generations

Max

Daily production

5,900/month

Get Started

Credits per month

150,000

≈2,700 images or ≈30 videos


At once

8 parallel generations

Ultimate

High-volume

9,900/month

Get Started

Credits per month

260,000

≈4,700 images or ≈55 videos


At once

10 parallel generations

Enterprise

Teams & agencies

1,00,000+/month

Contact Us

Credits per month

Custom volume

≈1,000 images or ≈12 videos


At once

Custom concurrency

  • Credits refresh monthly
  • Top-up additional credits anytime
  • Unused credits don't roll over

Not ready for the commitment?

Pay as you go

Free to sign up. Buy credits when you need them, same models, same features as every plan above.

Buy Credits

Credit packages

15 credits per 1

You pay You get
500 7,500 credits
1,000 15,000 credits
2,000 30,000 credits

FAQ

Frequently asked questions

Everything you need to know about Text to Speech on BeHooked.

Which voices can I choose from?

Forty-six named voices across three families. ElevenLabs has twenty: Aria, which is the default, plus Roger, Sarah, Laura, Charlie, George, Callum, River, Liam, Charlotte, Alice, Matilda, Will, Jessica, Eric, Chris, Brian, Daniel, Lily and Bill. Minimax has seventeen: Wise Woman, the default, plus Friendly Person, Inspirational Girl, Deep Voice Man, Calm Woman, Casual Guy, Lively Girl, Patient Man, Young Knight, Determined Man, Lovely Girl, Decent Boy, Imposing Manner, Elegant Man, Abbess, Sweet Girl 2 and Exuberant Girl. Qwen has nine: Vivian, the default, plus Serena, Uncle Fu, Dylan, Eric, Ryan, Aiden, Ono Anna and Sohee. Cartesia offers exactly one voice and no picker.

How many models are there, and which should I pick?

Nine. ElevenLabs v3 is the most expressive and reads emotion out of the writing; Multilingual v2 holds one voice across thirty languages; Turbo v2.5 is half the price and most of the quality, which makes it the draft option. Minimax Speech 2.8 Turbo gives you emotion, pitch and volume as real controls rather than hints, and 2.8 HD is the same voices at higher fidelity. Qwen3 TTS 1.7B covers eleven languages with automatic detection and 0.6B is the smaller, faster one with the same voices. Cartesia Sonic 3.5 finishes while you wait. Minimax Voice Clone reads your script in a voice from a recording you supply.

How is it priced?

Per 1000-character block, at the rate of the model you choose, with a one-block floor — a single short line still costs one block, so it is worth batching short lines into one run where the script allows. ElevenLabs Turbo v2.5 is the cheapest of the ElevenLabs three at about half the price. Failed runs are refunded.

Can I change the pace?

Yes, on every family. Speed runs from 0.7 to 1.2 on ElevenLabs, Qwen and Cartesia, and from 0.5 to 2 on Minimax. The narrow band is narrow deliberately: pushed further, the read stops sounding like a person. Before reaching for it, note that these models read punctuation as direction — a comma is a breath and a full stop is a beat, so write the pauses in rather than asking for them.

What does stability actually do?

It is an ElevenLabs control running from 0 to 1, defaulting to 0.5. Low values give a more expressive read that varies between takes; high values give a consistent, flatter one. Similarity, at 0.75 by default, is the separate control for how closely the render holds to the reference voice, and style, at 0 by default, pushes the delivery away from neutral. None of the three appears on the other families.

Can I make it sound happy, or angry?

On Minimax, yes — it takes one of seven named emotions: neutral, happy, sad, angry, fearful, disgusted or surprised, alongside a pitch shift of −12 to +12 semitones. Both are Minimax-only. On ElevenLabs the equivalent handle is the style control, from 0 to 1.

How long a script can I run at once?

Up to 5000 characters, enforced as you type. Longer scripts need splitting into runs — break at a paragraph rather than mid-sentence so the phrasing holds across the seam.

What languages does it cover?

Qwen3 TTS covers eleven with automatic detection: Auto, English, Chinese, Spanish, French, German, Italian, Japanese, Korean, Portuguese and Russian. ElevenLabs Multilingual v2 is the other multilingual option, holding one voice across thirty languages.

Paste a script and hear it read

Experience Text to Speech and every specialized AI app in your workspace.

Try Text to Speech