Stability AI released Stable Audio 3.0 today, a family of four generative audio models that create songs up to 6 minutes 20 seconds long. The large model more than doubles the maximum output length of Stable Audio 2.0, which shipped in 2024 and capped at roughly 3 minutes. The company is positioning the release as professional-grade music generation, built on fully licensed training data.
“The top model can generate professional-grade music of more than six minutes long”— Stability AI, Company statement
The four models span 459M parameters for the small and small SFX variants, 1.4B parameters for the medium model, and 2.7B parameters for the large model. The two small models generate audio up to 2 minutes and are designed for on-device deployment. The medium and large models both hit the 6-minute-20-second ceiling and maintain musical structure and melodic coherence across full compositions.
Stability AI is releasing the small SFX, small, and medium models with open weights for modification and commercial use under $1M annual revenue. The large model is available only via API or paid self-hosting, with an enterprise license required for companies exceeding the $1M threshold. This tiered approach mirrors the company's strategy with Stable Diffusion, where the most capable models remain commercially gated.
Key facts
- 01Stable Audio 3.0's large model generates songs up to 6 minutes 20 seconds long, more than double the 2024 version's capability.
- 02Four model sizes released: 459M-parameter small/SFX, 1.4B-parameter medium, and 2.7B-parameter large.
- 03Three models ship with open weights; the large model requires API access or paid self-hosting.
- 04Enterprise license required for companies exceeding $1M in revenue using the large model.
- 05Models built on fully licensed data following Stability AI's 2025 deals with Warner Music Group and Universal Music Group.
The licensing framework addresses the primary legal risk in AI music generation. Suno and Udio are both defending lawsuits from major labels over training-data provenance. Stability AI signed deals with Warner Music Group and Universal Music Group in 2025, securing licensed datasets for audio model development. The company confirmed that Stable Audio 3.0 was trained exclusively on that licensed corpus.
Stable Audio Open, released in 2024, generated audio up to 47 seconds long. The jump to 6-plus minutes represents a step change in usable output for full-track creation. Competitors including Google and ElevenLabs have shipped music-generation tools in the past year, but most cap output at 2–3 minutes or require stitching multiple generations.
Stability AI is building a professional musician product suite around the new models but did not disclose features or launch timelines. Ethan Kaplan, former chief digital officer at Universal Audio and Fender, joined the company to lead the professional music vertical. The hire follows a pattern across AI music startups hiring label and publisher executives to build industry credibility.
Suno hired former Merlin CEO Jeremy Sirota as chief commercial officer earlier this year. ElevenLabs brought in Derek Cournoyer from indie publisher Kobalt as a strategy lead for its music business. The executive moves signal that AI music companies are treating label relationships and licensing as competitive moats, not compliance overhead.
The large model's 6-minute ceiling is still shorter than most commercial tracks, which average 3–4 minutes but can run 7–10 minutes for album cuts. Extending coherent generation beyond 6 minutes without degrading musical structure remains an unsolved problem. Stability AI did not publish benchmark results comparing Stable Audio 3.0 to Suno v4 or Udio's latest models.
The medium model's open-weight release lowers the barrier for developers building music tools on licensed data. A 1.4B-parameter model generating 6-minute songs can run on consumer hardware, enabling local music creation without API costs. Whether that drives adoption depends on output quality relative to closed competitors.
Stability AI's pivot to licensed training data follows mounting legal pressure on generative AI companies. The Warner and Universal deals give the company a defensible position in music generation, but the terms of those licenses are undisclosed. If the labels secured revenue-sharing clauses or per-generation royalties, Stability AI's cost structure could differ materially from open-weight text models.
The audio model family marks Stability AI's first major product release under its post-restructuring strategy. The company has narrowed its focus to image, video, and audio generation after spinning out or shuttering earlier bets. Stable Audio 3.0 is the clearest test yet of whether open-weight models with enterprise upsells can compete against API-only players in generative media.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




