Skip to main content
An independent MiniMax Audio guide

MiniMax Audio: AI Voice, Speech and Music Generator

MiniMax Audio AI turns a script into natural speech, clones a voice from a short sample, designs new voices from a description, and writes music. This independent guide explains what MiniMax AI audio actually does, which model to pick, the output formats, what it costs, and how the API works.

Diagram in dark mode: MiniMax Audio's four capabilities - text to speech, voice cloning, voice design and music - and the speech-2.8-hd, speech-2.8-turbo, speech-2.6 and music-2.0 models behind them

OVERVIEW

What is MiniMax Audio?

MiniMax Audio is the audio product family from the AI lab MiniMax. One account and one API key cover text to speech, voice cloning, voice design and music generation. The current generation is the Speech 2.8 family, which added sound tags such as (laughs) and (sighs), clones a voice from roughly ten seconds of reference audio, and supports more than 40 languages with native-speaker-validated prosody.

  • Text to speech across 40+ languages
  • Voice cloning from a ten-second sample
  • Voice design written as a description
  • Music generation with vocals and backing

CAPABILITIES

What MiniMax Audio can do

Four capabilities, one account, one API key.

MiniMax text to speech

Turn a script into speech using 300+ system and cloned voices. Volume, pitch and speed are adjustable, voices can be mixed proportionally, and output can be streamed as it renders.

MiniMax voice cloning

Build a reusable voice from a sample of ten seconds to five minutes, supplied as mp3, m4a or wav under 20 MB. The clone keeps vocal texture, breathiness and pace, then behaves like any system voice.

MiniMax music

Generate songs from a text prompt with the Music 2.0 model, including vocals and instrumental backing, at sample rates up to 44.1 kHz.

MiniMax voice design

Describe a voice in words - age, tone, accent, delivery - and get a new synthetic voice without recording any reference audio at all.

SPECS

MiniMax Audio at a glance

The numbers that decide whether it fits your project.

40+

Languages supported

300+

System and cloned voices

~10 sec

Audio needed to clone a voice

HOW TO

How to use MiniMax Audio

Four steps from a blank page to a finished audio file.

  1. 01

    Pick the capability

    Choose text to speech, voice cloning, voice design or music. They share one account, so the choice only decides which screen or endpoint you use.

  2. 02

    Choose a voice and a model

    Take a system voice or a clone you created, then select speech-2.8-hd for fidelity or speech-2.8-turbo for low latency.

  3. 03

    Write the script with sound tags

    Paste the text and drop inline cues such as (laughs) or (sighs) where delivery should change. Adjust volume, pitch and speed if the default read is off.

  4. 04

    Render and export

    Generate, listen, and export as mp3, wav or flac. For streaming playback the API returns mp3 as it renders.

MODELS

MiniMax audio model versions

Match the MiniMax audio model to your quality, latency and budget.

MiniMax Speech 2.8 HD

The high-fidelity tier, for audiobooks, narration and anything published once and replayed often.

MiniMax Speech 2.8 Turbo

The low-latency tier, built for voice agents, live assistants and streaming playback.

MiniMax Speech 2.6

The previous generation, still callable when you need output consistent with earlier renders.

MiniMax Music 2.0

The music model: vocals plus instrumental arrangement generated from a text prompt.

Voice Design

Builds a voice from a written description instead of a reference recording.

Sound Tags

Inline cues such as (laughs) and (sighs) that steer delivery from inside the script. Introduced with Speech 2.8.

MiniMax Audio logo

The MiniMax audio API and audio formats

The MiniMax audio format you receive depends on how you call it: the synchronous text-to-audio (T2A) endpoint returns mp3, wav or flac, while streaming requests return mp3 only. Music output is mp3, wav or pcm at up to 44.1 kHz and 256 kbps. A single MiniMax audio API key covers speech, cloning, voice design and music.

FIT

Who it suits, and who it does not

An honest read on where this tool wins and where it costs you time.

It fits you if price per character is the constraint

The clearest advantage here is cost at volume. If you are rendering long-form narration, a large library of product descriptions, or a podcast back-catalogue, the per-character rate is what decides your monthly bill, and this sits well below the premium Western vendors. Sound tags help too: being able to write a laugh or a sigh into the script means fewer re-renders chasing a delivery you could not otherwise control.

It fits you if you need many languages without many vendors

Coverage of 40+ languages with prosody checked by native speakers matters most when you are localising one script into a dozen markets. Doing that across several single-language vendors means several contracts, several quality baselines, and several integrations. One endpoint and one key removes that overhead, and a cloned voice can carry across languages rather than being re-cast per market.

It fits you less if you need a large curated voice library

If your work depends on browsing hundreds of professionally directed, rights-cleared character voices and picking the exact right one, the specialist vendors still have deeper catalogues and better discovery tooling. The answer here is usually to design or clone the voice you need rather than to shop for it, which is a different workflow and takes longer the first time.

Check these before you commit

Three things change without announcement and will not appear in any guide fast enough: the exact model IDs available on your account route, the current per-character rate, and the free allowance. Pin a model version in your integration rather than tracking the latest tag, set a spend alert at the account level because all four capabilities draw on one balance, and read the commercial terms for your plan before you publish anything you sell.

MiniMax Audio pricing

MiniMax bills audio pay-as-you-go. Figures below are MiniMax's published list rates, last checked September 2026 - confirm them on MiniMax's own pricing page before you budget.

speech-2.8-turbo

$60

per 1M characters

The low-latency text to speech tier.

speech-2.8-hd

$100

per 1M characters

The high-fidelity text to speech tier.

Rapid voice cloning

$1.50

per voice

Charged once when the voice is created, not per generation.

Voice design

$3.00

per voice

Charged once for each designed voice.

GO DEEPER

MiniMax Audio resources

The questions people ask most, and where each one is answered.

Still deciding on MiniMax Audio?

Start with what it does and what it costs, then try it on MiniMax's own site.

FAQs

MiniMax Audio questions, answered

Newsletter

Join the community

Subscribe to our newsletter for the latest news and updates