Skip to main content
All posts
Models

MiniMax Speech 2.8 HD vs Turbo - which model should you use?

The practical difference between speech-2.8-hd and speech-2.8-turbo, and how to pick one without burning credits on the wrong tier.

MiniMax ships the Speech 2.8 family in two tiers, and the names undersell how different the decision is. Picking wrong does not break anything — it just costs you either money or milliseconds, forever, on every request.

The short version

Use speech-2.8-hd when the audio is rendered once and listened to many times. Use speech-2.8-turbo when someone is waiting for the audio to start.

speech-2.8-hdspeech-2.8-turbo
List price$100 / 1M characters$60 / 1M characters
Optimised forfidelitylatency
Typical useaudiobooks, narration, adsvoice agents, live assistants

List prices checked September 2026 on MiniMax's pay-as-you-go page. Confirm before you budget — rates vary by account route.

Why the split exists

Both tiers share the Speech 2.8 improvements: sound tags such as (laughs) and (sighs) that steer delivery from inside the script, a re-engineered pipeline that cuts background noise, and prosody validated by native speakers across 40+ languages.

What differs is how much work the model does per token before it hands you audio. HD spends more of it, and you hear that in sustained vowels, breath, and the tail of a sentence. Turbo spends less and starts returning audio sooner.

How to actually decide

Ask one question: does a human notice the wait?

  • A user clicks "play" on an article you generated last night → nobody is waiting. Use HD.
  • A support bot has to answer in a phone call → every 200ms is visible. Use Turbo.
  • You are batch-rendering 400 product descriptions overnight → nobody is waiting. Use HD, unless the character volume makes the price difference matter more than the quality difference.

That last case is the only genuinely hard one. At 40% cheaper, Turbo wins on any workload where the character count is large and the listener is casual.

Test it on your own script

Model comparisons on a marketing page are close to useless, because quality differences cluster around specific things — proper nouns, numbers, questions, emotional lines. Take 200 words of your actual content, render it through both, and listen on the device your audience will use. The difference will either be obvious or it will not, and either answer settles it.

What about Speech 2.6?

2.6 is still callable. The reason to use it is consistency: if you shipped a library of audio on 2.6 and need new files to match, staying on 2.6 beats re-rendering everything. For new work, there is no quality argument for it.