MiniMax voice cloning - what you need, what it costs, and what you must not do
The upload requirements for MiniMax rapid voice cloning, how the $1.50-per-voice billing works, and the consent rule that matters more than any of it.
MiniMax's rapid voice cloning turns a short recording into a reusable voice. Once created, it behaves like any system voice — you call it by ID and it costs the same per character as any other render.
The upload requirements
- Format: mp3, m4a, or wav
- Length: 10 seconds to 5 minutes
- Size: 20 MB maximum
Roughly ten seconds is genuinely enough. What the model captures from that sample is vocal texture, breathiness, and speaking pace — not vocabulary or accent quirks that only show up over minutes.
Clean beats long
The single biggest lever on clone quality is the noise floor of your source, not its duration. Thirty seconds recorded on a decent microphone in a quiet room produces a noticeably better clone than four minutes recorded in a café. If you have a choice, record fresh rather than pulling from a podcast or a video call.
Practical checklist for the sample:
- One speaker, no overlapping voices, no music bed.
- Normal speaking register — not a performance, not a whisper.
- A few full sentences with varied punctuation, so pacing has something to learn from.
- No compression artefacts. Re-encoding a low-bitrate mp3 makes it worse.
How billing works
Rapid voice cloning is charged once per voice created — $1.50 at MiniMax's published list rate — not per generation. After that the voice is called like a system voice at normal text-to-speech rates. Voice Design, which builds a voice from a written description with no reference audio at all, is $3.00 per voice on the same basis.
Prices checked September 2026. Confirm on MiniMax's pricing page before budgeting.
The practical consequence: iterating on a clone is not free. Each attempt is a new voice and a new charge. Get the source recording right before you upload rather than cloning five times and picking a favourite.
The part that is not optional
Do not clone a real person's voice without that person's clear, documented consent for the specific use you intend.
This is not a style guideline. Voice cloning without consent exposes you to liability under likeness, publicity, and in several jurisdictions specific synthetic-media laws — and the person whose voice it is has no way to withdraw it once your audio is published. "They are a public figure" is not consent. "It is only for internal use" is not consent either, because internal audio leaks.
If you need a voice that sounds like a particular kind of person rather than a particular person, use Voice Design instead. Describing the age, tone, accent, and delivery you want produces a synthetic voice that belongs to nobody — which is the entire point.