Scripts read back in a cloned voice
A cloned voice has no thumbnail. Each entry is the text that was sent with the reference and the waveform of what came back.
- Prompt
“We built this because we kept re-recording the same forty seconds every time the script moved.”
0:07 - Prompt
“In this lesson we will look at what the model does when the reference recording is short.”
0:06 - Prompt
“That is not what I expected to happen at all.”
0:03 - Prompt
“Everything you generate stays in your workspace, and you can come back to any of it later.”
0:08
One recording, one script, one request
Upload a reference recording
Five to thirty seconds of clean speech, from a voice you have the rights to use. Minimax requires a sample. Cartesia's is optional — leave it off and it falls back to its own default voice.
Write the script
The text goes in alongside the sample. Cloning and speaking are one request on Minimax rather than two separate jobs, so on this surface there is no saved voice to select first.
Set the engine and the read
Four Minimax engines at four price points, plus speed, pitch and emotion. The finished audio lands in your workspace library, and failed runs are refunded.
What Voice Clone takes, and what it returns
What Voice Clone needs from you
- Reference audio
- Five to thirty seconds of clean speech — required on Minimax, optional on Cartesia
- Script
- The text to be spoken, sent in the same request
- Price
- Per 1000 characters, by engine: 301 on 02 HD, 181 on 02 Turbo, 271 on 01 HD, 91 on 01 Turbo
- Rights
- Permission from the person whose voice you are cloning
Export specifications
- Output
- Your script, read in the cloned voice
- Steps
- One request on Minimax — cloning and speech together
- Price
- Per 1000 characters, from 91 credits on 01 Turbo to 301 on 02 HD
- Failed runs
- Refunded
- Delivery
- Rendered audio file in your workspace library
The sample, the engine, and how the line is read
Minimax takes the sample and the script together in one request. The engine you choose sets both the fidelity and the price.
-
Clone audio
One reference recording, five to thirty seconds
The sample the voice is built from, and clean speech is what it needs. Required on Minimax; optional on Cartesia, which falls back to its own default voice when you leave it off.
-
Engine
Minimax Speech 02 HD, 02 Turbo, 01 HD or 01 Turbo
Four engines at 301, 181, 271 and 91 credits per 1000 characters respectively. It defaults to 02 HD, the dearest and the highest fidelity — 01 Turbo is a third of that price if the read is a draft.
-
Script
The text to speak
Cloning and generation happen in one step on this surface, so there is no voice to save and select first. Story Mode is the surface that does keep a voice library.
-
Speed
0.5–2, default 1
The Minimax range, from a slow read to a compressed one.
-
Pitch
−12 to +12, default 0
Semitone shift on the cloned voice, without changing the pace.
-
Emotion
Neutral, happy, sad, angry, fearful, disgusted, surprised
The same seven Minimax emotions available on Text to Speech, applied to the cloned voice.
Engine choice
Pick the one that matches your source.
-
Minimax Speech 02 HD
DefaultThe default engine and the highest fidelity of the four. 301 credits per 1000 characters.
A finished read
-
Minimax Speech 02 Turbo
FasterThe faster 02 engine, at a little over half the price of HD. 181 credits per 1000 characters.
Iterating on a long script
-
Minimax Speech 01 HD
Previous HDThe previous generation at high fidelity. 271 credits per 1000 characters.
Matching earlier work
-
Minimax Speech 01 Turbo
CheapestThe cheapest of the four by some distance. 91 credits per 1000 characters.
Drafts and volume
-
Cartesia Sonic 3.5
InstantInstant cloning, and the one engine whose sample is optional — without one it reads in its own default voice.
Speed on a clean reference
Built for people who need their own voice at scale
-
Creators
Keep your own voice on a channel without recording every line, once you have five to thirty seconds of clean speech.
-
Course authors
Update a module's narration without matching a microphone and a room you no longer have.
-
Brands with a signature voice
Keep one talent consistent across campaigns, with that talent's agreement in place.
Built for Growth at Every Stage
Every plan includes every model and every feature. Plans only change how many credits you get and how many generations run at once.
- Credits refresh monthly
- Top-up additional credits anytime
- Unused credits don't roll over
Not ready for the commitment?
FAQ
Frequently asked questions
Everything you need to know about Voice Clone on BeHooked.
Do I need permission to clone a voice?
Yes. You need the rights to the voice you are cloning — your own, or one whose owner has agreed to this use. Do not upload a recording of someone who has not.
How long should the reference recording be?
Five to thirty seconds of clean speech. Minimax requires a sample and will not run without one. Cartesia Sonic 3.5 is the exception: its sample is optional, and without one it simply reads in its own default voice.
Which engine should I choose, and what does it cost?
There are four Minimax engines, priced per 1000 characters: Speech 02 HD at 301 credits, 02 Turbo at 181, 01 HD at 271 and 01 Turbo at 91. It defaults to 02 HD, which is the dearest and the highest fidelity — if the read is a draft, 01 Turbo is under a third of the price for the same script. Failed runs are refunded.
Is cloning a separate step from generating speech?
Not on this surface. The reference recording and the script go in the same request on Minimax, and you get the script back in that voice, so there is no voice to save and select first. Story Mode works differently: it keeps a persistent voice library, where a voice you give a character is kept and offered on the next film.
Can I set the emotion, pitch or pace of a cloned voice?
Yes. The cloned voice takes the same Minimax controls as a stock one: speed from 0.5 to 2, pitch from −12 to +12 semitones, and one of seven named emotions — neutral, happy, sad, angry, fearful, disgusted or surprised.
Where does it open in the dashboard?
On the Text to Speech surface. Voice Clone is the same workspace with a reference recording attached, rather than a separate app.