Stability AI's Stable Audio 3.0 generates music and sound up to about six minutes from text, then lets you edit, extend, and reshape it. Trained entirely on licensed data.
Small SFX and Small (459M), Medium (1.4B), and the flagship Large (2.7B) span on-device sound design to full studio compositions.
Medium and Large hold musical structure and melodic tone across compositions running roughly six minutes and twenty seconds.
Inpainting supports single-segment and multi-segment editing plus causal continuation to extend an existing track past its endpoint.
Trained on fully licensed data; under the Community License you own your outputs and can distribute and commercialize them.
Stable Audio 3.0 is Stability AI's May 2026 audio generation release, structured as a family of four models rather than a single monolith. The Small SFX and Small models (459M parameters each) target on-device sound effects and short-form music up to roughly two minutes, while the Medium (1.4B) and Large (2.7B) models generate full compositions running up to about six minutes and twenty seconds, holding musical structure and melodic tone across that span. Everything is trained on fully licensed data, and Stability positions the family as a foundation the wider audio community can build on.
Beyond text-to-audio, the models are designed for iteration. Inpainting covers single-segment editing, multi-segment editing, and causal continuation, so creators can rework a passage or extend a track past its original endpoint instead of regenerating from scratch. LoRA fine-tuning lets teams adapt the models to a specific sound. The Small SFX, Small, and Medium tiers ship as open weights on Hugging Face under the Stability AI Community License, while the flagship Large model is available through the Stability AI API and paid self-hosting, with an enterprise license for organizations above roughly $1M in annual revenue.
The Medium and Large models produce full compositions of roughly six minutes and twenty seconds, maintaining musical structure and melodic tone rather than looping short clips.
Small SFX and Small (459M) handle sound effects and short music suited to consumer hardware; Medium (1.4B) scales to full tracks; Large (2.7B) is the flagship quality tier.
Edit a single segment, revise multiple sections at once, or continue a track past its endpoint with causal continuation, keeping the parts you already like.
Small SFX, Small, and Medium are released as open-weight models on Hugging Face, so developers can run and modify them locally.
The family is trained on fully licensed data, a deliberate stance on rights that pairs with output ownership under the Community License.
Teams can adapt the models to a specific instrument, genre, or sonic signature through LoRA fine-tuning instead of prompting alone.
Under the Stability AI Community License, users own their generated audio and can distribute and commercialize it, with enterprise licensing for larger organizations.
| Small SFX | 459M params, sound effects, on-device |
|---|---|
| Small | 459M params, music up to ~2 min |
| Medium | 1.4B params, up to ~6:20 |
| Large | 2.7B params, flagship, API / self-host |
| Input | Text prompt (plus existing audio for editing) |
|---|---|
| Output | Music and sound effects |
| Max duration | ~2 min (Small) / ~6 min 20 sec (Medium, Large) |
| Editing | Single- & multi-segment inpainting, causal continuation |
| Sample rate | 44.1kHz stereo (Stable Audio family standard) |
| Open weights | Small SFX, Small, Medium (Hugging Face) |
|---|---|
| Large access | Stability AI API + paid self-hosting |
| License | Stability AI Community License |
| Enterprise | Required above ~$1M annual revenue |
| Fine-tuning | LoRA supported |
Generate score-length tracks that hold structure across a full scene, then trim or extend sections with continuation to match your edit.
The lightweight Small SFX model targets on-device generation of sound effects and short cues suited to consumer hardware.
Draft ideas up to roughly six minutes, then use inpainting to revise a bridge or chorus without regenerating the whole piece.
Fine-tune with LoRA on a signature palette so a studio or brand can produce consistent, on-style audio at scale.
Because the family is trained on licensed data and outputs are ownable under the Community License, it suits teams wary of training-data provenance.
| Feature | Stable Audio 2.5 | Stable Audio 3.0NEW |
|---|---|---|
| September 2025 | May 20, 2026 | |
| Single production-focused model | Four-tier family (459M-2.7B) | |
| No (API / cloud) | Yes for Small SFX, Small, Medium | |
| Multi-part composition (intro/development/outro) | Inpainting + multi-segment editing + continuation | |
| Shorter-form compositions | Up to ~6 min 20 sec (Medium/Large) |
Only the Medium and Large models reach the ~6:20 mark; the Small tiers cap around two minutes, so long-form work needs the larger models.
The 2.7B Large model is API-only plus paid self-hosting, not a downloadable open-weight release like the smaller three tiers.
Organizations above roughly $1M in annual revenue need an enterprise license, so commercial use at scale carries additional terms.
As with any model, coherence and mix quality can vary by prompt and genre, so outputs benefit from human listening and cleanup before release.
We're bringing Stability AI's licensed-data audio family, with six-minute tracks and segment-level editing, to SharkFoto as an upgrade to our existing Stable Audio 2.5 support. Explore SharkFoto's available AI tools today and watch this page for launch.
Try Stable Audio 3.0 now