On August 14, MiniMax officially launched its next-generation music generation model MiniMax-Music3. Users can generate complete songs up to 5 minutes long by inputting lyrics and a music description, with output in 32kHz, 16-bit stereo WAV format. The model weights are fully open-sourced on Hugging Face.

Technical Architecture: 8B Global Model for Structure, 0.6B Local Model for Detail
The capability is powered by a hierarchical autoregressive architecture with two complementary models.
The Global LLM (8B parameters), initialized from Qwen3-8B, predicts the first RVQ codebook frame by frame, modeling the song's long-range semantics and structural progression—intro, verse, chorus, bridge, and outro. The Local LLM (0.6B parameters) predicts the remaining acoustic codebooks within each frame, restoring fine-grained acoustic details.
Training proceeds in two stages: first aligning the Global LLM's embedding and output layers to musical semantic tokens, then jointly training both models on all codebooks. On the output side, the model uses a 2.4B Flow Matching module and a 123M Flow-VAE decoder to synthesize continuous audio from the language model's hidden states, rather than decoding directly from discrete tokens, resulting in clearer vocal articulation and more coherent instrumental texture.
Fine-Grained Control: From Prompts to Arrangement Instructions
Music3.0's core upgrade lies in its ability to understand creative intent. The model focuses on the hardest aspects to capture with a simple prompt: understanding the creator's expressive intent, sustaining it across a complete song, rendering instruments with clarity and physical realism, and generating vocals that sound performed rather than synthesized.
To this end, MiniMax introduced a Structured Captions framework. Users can precisely control generation through three sections: Global Metadata (genre, BPM, key, scale, emotional progression), Vocal Details (gender, timbre, performance style, harmony, effects), and Arrangement (instruments, section-level evolution, groove, bass, percussion, texture, spatial effects).
Lyrics input supports explicit section tags: [Intro], [Verse], [Pre-Chorus], [Chorus], [Bridge], [Instrumental], [Solo], [Outro], allowing precise specification of song structure.
Open Source and Pricing: $0.15 Per Song, Self-Hostable
Music3.0 is fully open-sourced, with model weights available for download on Hugging Face and support for self-hosting at an estimated ~22GB VRAM.
Official API pricing is $0.15 per song (up to 5 minutes), with a rate-limited free tier. By comparison, Suno v4 costs roughly 5 credits per song (Pro plan $10/month for 2,500 credits), and generations are typically only 2-4 minutes. MiniMax-Music3 outputs a 32kHz WAV file that users fully own, while competitors like Udio impose download restrictions.
Music3.0's "layered architecture + open-source strategy" precisely addresses two core pain points in AI music creation: long-range structural stability and creative controllability. While most music models remain in the "prompt lottery" stage, MiniMax transforms AI music generation from a game of chance into a programmable craft through structured descriptions. The combination of open weights and $0.15/song pricing puts this capability firmly in the hands of developers and creators.