Skip to main content
All posts
API

MiniMax audio formats and the T2A API, explained

Which formats the synchronous and streaming endpoints return, what the music model outputs, and how the pieces fit behind one API key.

The format question comes up early because it constrains everything downstream — what you can stream to a browser, what you can hand to an editor, and what you have to transcode.

Text to audio (T2A)

  • Non-streaming requests return mp3, wav, or flac.
  • Streaming requests return mp3 only.

That asymmetry is the thing to design around. If your player needs to start before the render finishes, you are getting mp3 and there is no negotiation. If you are writing files to storage for later playback, take wav or flac and keep the headroom for whatever post-processing comes next.

Music

The music model outputs mp3, wav, or pcm, at sample rates up to 44.1 kHz and bitrates up to 256 kbps. PCM is the useful one if the audio is going straight into a DAW or a mixing step — no decode, no generation loss.

One key, four capabilities

Text to speech, voice cloning, voice design, and music all sit behind the same MiniMax account and the same API key. In practice this means:

  • You do not provision separate credentials per capability.
  • A voice you clone through the API is immediately callable from a T2A request by ID.
  • Spend across all four draws on the same balance, so a runaway music job can starve your speech traffic. Set an alert on the account, not per endpoint.

Choosing a format in practice

SituationTake
Streaming to a web playermp3 (no choice)
Archive master you will re-cut laterwav or flac
Feeding a video editorwav
Music heading into a DAWpcm
Serving a mobile app over cellularmp3

Before you build

Formats, model IDs, and limits move. Everything here was checked in September 2026 against MiniMax's public API documentation — read the current docs before you commit an integration to them, particularly if you are pinning a model version.