Stability · Licensed-data audio generation, open where it counts.

Stable Audio 3.0

Open-weight AI audio, built for artists

Stability AI's Stable Audio 3.0 generates music and sound up to about six minutes from text, then lets you edit, extend, and reshape it. Trained entirely on licensed data.

~6:20
Max track length
4
Model tiers
459M-2.7B
Parameter range
3 of 4
Open-weight tiers
What's new

Open-weight AI audio, built for artists

01

Four-tier family

Small SFX and Small (459M), Medium (1.4B), and the flagship Large (2.7B) span on-device sound design to full studio compositions.

02

Tracks to ~6:20

Medium and Large hold musical structure and melodic tone across compositions running roughly six minutes and twenty seconds.

03

Edit, not just generate

Inpainting supports single-segment and multi-segment editing plus causal continuation to extend an existing track past its endpoint.

04

Licensed and ownable

Trained on fully licensed data; under the Community License you own your outputs and can distribute and commercialize them.

Overview

Model at a glance

Stable Audio 3.0 is Stability AI's May 2026 audio generation release, structured as a family of four models rather than a single monolith. The Small SFX and Small models (459M parameters each) target on-device sound effects and short-form music up to roughly two minutes, while the Medium (1.4B) and Large (2.7B) models generate full compositions running up to about six minutes and twenty seconds, holding musical structure and melodic tone across that span. Everything is trained on fully licensed data, and Stability positions the family as a foundation the wider audio community can build on.

Beyond text-to-audio, the models are designed for iteration. Inpainting covers single-segment editing, multi-segment editing, and causal continuation, so creators can rework a passage or extend a track past its original endpoint instead of regenerating from scratch. LoRA fine-tuning lets teams adapt the models to a specific sound. The Small SFX, Small, and Medium tiers ship as open weights on Hugging Face under the Stability AI Community License, while the flagship Large model is available through the Stability AI API and paid self-hosting, with an enterprise license for organizations above roughly $1M in annual revenue.

Developer
Stability AI
Announced
May 20, 2026
Modality
Text-to-audio, music & SFX
Model tiers
Small SFX, Small, Medium, Large
Parameters
459M / 459M / 1.4B / 2.7B
Max duration
Up to ~6 min 20 sec (Medium/Large)
Editing
Inpainting + causal continuation
Open weights
Small SFX, Small, Medium (Hugging Face)
Capabilities

Key features

Long-form generation to ~6:20

The Medium and Large models produce full compositions of roughly six minutes and twenty seconds, maintaining musical structure and melodic tone rather than looping short clips.

Four purpose-built tiers

Small SFX and Small (459M) handle sound effects and short music suited to consumer hardware; Medium (1.4B) scales to full tracks; Large (2.7B) is the flagship quality tier.

Inpainting and segment editing

Edit a single segment, revise multiple sections at once, or continue a track past its endpoint with causal continuation, keeping the parts you already like.

Open weights for three tiers

Small SFX, Small, and Medium are released as open-weight models on Hugging Face, so developers can run and modify them locally.

Fully licensed training data

The family is trained on fully licensed data, a deliberate stance on rights that pairs with output ownership under the Community License.

LoRA fine-tuning

Teams can adapt the models to a specific instrument, genre, or sonic signature through LoRA fine-tuning instead of prompting alone.

Output ownership

Under the Stability AI Community License, users own their generated audio and can distribute and commercialize it, with enterprise licensing for larger organizations.

Specs

Technical specifications

Model family

Small SFX459M params, sound effects, on-device
Small459M params, music up to ~2 min
Medium1.4B params, up to ~6:20
Large2.7B params, flagship, API / self-host

Generation & editing

InputText prompt (plus existing audio for editing)
OutputMusic and sound effects
Max duration~2 min (Small) / ~6 min 20 sec (Medium, Large)
EditingSingle- & multi-segment inpainting, causal continuation
Sample rate44.1kHz stereo (Stable Audio family standard)

Access & licensing

Open weightsSmall SFX, Small, Medium (Hugging Face)
Large accessStability AI API + paid self-hosting
LicenseStability AI Community License
EnterpriseRequired above ~$1M annual revenue
Fine-tuningLoRA supported
In practice

Use cases

Background music for video

Generate score-length tracks that hold structure across a full scene, then trim or extend sections with continuation to match your edit.

Game and app sound design

The lightweight Small SFX model targets on-device generation of sound effects and short cues suited to consumer hardware.

Music sketching and demos

Draft ideas up to roughly six minutes, then use inpainting to revise a bridge or chorus without regenerating the whole piece.

Custom-tuned sound libraries

Fine-tune with LoRA on a signature palette so a studio or brand can produce consistent, on-style audio at scale.

Rights-conscious commercial audio

Because the family is trained on licensed data and outputs are ownable under the Community License, it suits teams wary of training-data provenance.

Generational leap

FeatureStable Audio 2.5Stable Audio 3.0NEW
September 2025May 20, 2026
Single production-focused modelFour-tier family (459M-2.7B)
No (API / cloud)Yes for Small SFX, Small, Medium
Multi-part composition (intro/development/outro)Inpainting + multi-segment editing + continuation
Shorter-form compositionsUp to ~6 min 20 sec (Medium/Large)
Honest look

Current limitations

Best length varies by tier

Only the Medium and Large models reach the ~6:20 mark; the Small tiers cap around two minutes, so long-form work needs the larger models.

Flagship is not open weight

The 2.7B Large model is API-only plus paid self-hosting, not a downloadable open-weight release like the smaller three tiers.

Enterprise licensing threshold

Organizations above roughly $1M in annual revenue need an enterprise license, so commercial use at scale carries additional terms.

Generative audio still needs review

As with any model, coherence and mix quality can vary by prompt and genre, so outputs benefit from human listening and cleanup before release.

FAQ

Frequently asked questions

What is Stable Audio 3.0?
It is Stability AI's May 2026 audio generation family that turns text prompts into music and sound effects, with tracks up to about six minutes and built-in editing. It comes as four tiers: Small SFX, Small, Medium, and Large.
How long can Stable Audio 3.0 tracks be?
The Small models generate up to about two minutes, while the Medium (1.4B) and Large (2.7B) models produce full compositions of roughly six minutes and twenty seconds while maintaining musical structure.
Which models are open weight?
Small SFX, Small, and Medium are released as open-weight models on Hugging Face under the Stability AI Community License. The flagship Large model is available through the Stability AI API and paid self-hosting.
Can I edit audio, not just generate it?
Yes. Inpainting supports single-segment editing, multi-segment editing, and causal continuation, so you can revise part of a track or extend it past its original endpoint.
Is the training data licensed?
Stability AI states the family is trained on fully licensed data, and under the Community License users own their outputs and can distribute and commercialize them.
How is it different from Stable Audio 2.5?
Version 3.0 moves from a single model to a four-tier family, adds open weights for three tiers, extends length to about six minutes, and introduces multi-segment inpainting and causal continuation.
Can I use Stable Audio 3.0 on SharkFoto?
Not yet. Stable Audio 3.0 is coming soon to SharkFoto as an upgrade to the Stable Audio 2.5 support already offered. In the meantime you can explore SharkFoto's currently available AI creative tools for image and video.

Stable Audio 3.0 is coming soon to SharkFoto

We're bringing Stability AI's licensed-data audio family, with six-minute tracks and segment-level editing, to SharkFoto as an upgrade to our existing Stable Audio 2.5 support. Explore SharkFoto's available AI tools today and watch this page for launch.

Try Stable Audio 3.0 now