MiniMax Music 3.0: Next-Generation Open-Weights, Production-Ready & Versatile Music Model - MiniMax Research
The model composes, arranges, performs, and produces a complete song of up to five minutes in a single generation.
- An 8B global LLM initialized from Qwen3.5-8B and a 0.6B local LLM jointly model song structure and acoustic detail.
- Structured Captions describe emotion, instrumentation, and vocal delivery at section level; a Prompt Enhancement System expands simple user prompts.
- Audio is rendered by fusing hidden states from both LLMs into a 2.4B flow-matching module and a 123M Flow-VAE.
- Lyrics can include section tags such as [intro], [verse], [chorus], and [bridge] to define the song's macrostructure.
Music creators can now generate complete, production-ready songs up to five minutes from a simple prompt, with coherent structure and natural vocals.
Sources
Read this as text
Back to the AI news