FLUX 3 is Black Forest Labs' unified multimodal frontier model — image, video, audio and action from a single set of weights. Its first generally available product, FLUX 3 Video, generates clips of up to 20 seconds with optional native audio from text, images or video.
Unlike FLUX.2, FLUX 3 learns jointly from images, video and audio in one architecture, then extends to action prediction. FLUX 3 Video draws on that shared understanding of scenes, motion and sound.
FLUX 3 Video generates a single clip of up to 20 seconds, and through the BFL API it can also extend existing videos or transition between keyframes.
Switch on native audio to generate sound together with the video in one pass — BFL doesn't charge extra for audio.
FLUX 3 is built on Self-Flow, BFL's flow-matching approach that aligns generation and understanding across modalities inside the same model.
FLUX 3 is the third-generation flagship from Black Forest Labs (BFL), the Freiburg, Germany lab behind the FLUX family. Announced on 23 July 2026, it is a departure from its predecessors: rather than a standalone image model, FLUX 3 is a unified multimodal frontier model that jointly learns from images, video and audio within a single architecture and can be extended to predict actions. BFL frames it as a 'foundation layer for visual intelligence' — one network spanning image, video, audio and physical-AI action prediction, delivered across product lines — FLUX 3 Video and FLUX 3 Action today, with FLUX 3 Image and the open-weight FLUX 3 Dev on BFL's roadmap.
This page focuses on FLUX 3 Video, the first FLUX 3 product to reach general availability: it opened on the BFL API and select partner platforms on 4 August 2026. It covers text-to-video, image-to-video, video-to-video editing, video extension and keyframe transitions, generates up to 20 seconds per clip with optional native audio at no extra cost, and renders in HD, FHD, QHD and UHD-4K tiers. FLUX 3 Action, a 7B open-weight model for robotics, followed on 23 September 2026. FLUX 3 Image, for image generation and editing, has been announced but not yet released, with no public API as of September 2026.
FLUX 3 is trained jointly across video, images and audio, so FLUX 3 Video draws on a shared model of scenes, motion and sound. BFL reports stronger complex-prompt handling in its preliminary mid-training evaluations.
Generate from a text prompt, animate a still image, or edit an existing clip with video-to-video. FLUX 3 Video also supports extending videos and keyframe transitions.
FLUX 3 Video can add native audio to a clip in the same generation, so sound and picture come from one model instead of a separate audio step.
FLUX 3 is built on Self-Flow, BFL's approach that combines a flow-matching objective with self-supervised feature reconstruction to align generation and understanding inside a single model. Self-Flow was first introduced by BFL in March 2026.
FLUX 3 Video renders in HD, FHD, QHD and UHD-4K tiers based on pixels per frame, with pricing on the BFL API that scales by tier.
Video and action ship first as product lines over one backbone, with FLUX 3 Image for image generation and editing and the open-weight FLUX 3 Dev on BFL's roadmap; FLUX 3 Dev is planned for later in 2026.
| Name | FLUX 3 Video |
|---|---|
| Developer | Black Forest Labs (Freiburg, Germany) |
| Announced | 23 July 2026 |
| Model type | Unified multimodal flow model |
| Foundation | Self-Flow (flow matching + self-supervised feature reconstruction) |
| Modalities in one model | Image, video, audio, action |
|---|---|
| Generation modes | Text-to-video, image-to-video, video-to-video editing, extension, keyframe transitions |
| Max clip length | 20 seconds per generation |
| Audio | Optional native audio (no extra charge) |
| Resolution tiers | HD, FHD, QHD, UHD-4K (BFL API) |
| FLUX 3 Video | Generally available via the BFL API and partners since Aug 4, 2026 |
|---|---|
| FLUX 3 Action | 7B open weights released Sep 23, 2026 (FLUX Kommunity License) |
| FLUX 3 Image | Announced; not yet released (no public API as of September 2026) |
| FLUX 3 Dev (open weights) | Planned later in 2026 |
| SharkFoto | Text to Video & Image to Video: 720p/1080p, 5–20 s, native audio on/off |
Generate clips of up to 20 seconds with native sound for Reels, Shorts and TikTok from a single prompt, in vertical or landscape formats.
Animate a product photo into a lifestyle clip with image-to-video, using a start frame and an optional end frame to control how the shot opens and closes.
Turn storyboards and key frames into moving previews to test pacing, camera moves and mood before committing to a full production.
Produce short ads and promos with native audio generated in the same pass, reducing the need for a separate sound step.
Through the BFL API, FLUX 3 Video's video-to-video editing, extension and keyframe transitions can restyle or lengthen existing clips.
FLUX 3 is a fundamental shift from FLUX.2: where FLUX.2 is an image-only model family (FLUX.2 [dev] has 32B parameters), FLUX 3 is a unified multimodal backbone whose first release is FLUX 3 Video. FLUX 3 Image is still on BFL's roadmap, so no image-mode specs have been published yet.
| Feature | FLUX.2 | FLUX 3NEW |
|---|---|---|
| Scope | Image generation only | Unified image, video, audio and action |
| Architecture | 32B rectified-flow transformer + Mistral 3 VLM | Self-Flow multimodal flow model |
| Resolution | Up to 4MP (e.g. 2048x2048) | Video: HD to UHD-4K tiers |
| Video & audio | Not supported | Up to 20 s video with optional native audio |
| Availability | Generally available (Pro/Flex/Dev) | Video GA since Aug 2026; Image on the roadmap |
FLUX 3 Image, for image generation and editing, is on BFL's roadmap but had no public release or API as of September 2026; today FLUX 3 is available as FLUX 3 Video.
On SharkFoto, FLUX 3 renders 5–20 second clips at 720p or 1080p in Text to Video and Image to Video. The BFL API's QHD and 4K tiers, video-to-video editing, extension and keyframe transitions aren't offered here yet.
The open-weight FLUX 3 Dev backbone for content creation is still pending; only the robotics-focused FLUX 3 Action weights are public so far.
FLUX 3 is trained jointly across video, images and audio; teams that need a dedicated, production-proven image pipeline today may find FLUX.2 more predictable.
Turn a prompt or a still image into a 5–20 second clip at up to 1080p with optional native audio — FLUX 3 Video is available now in SharkFoto's Text to Video and Image to Video tools.
Try FLUX 3 now