Pricing Blog
Black Forest Labs · 20-second video with native audio from one multimodal backbone.

FLUX 3

One model for all visual intelligence

FLUX 3 is Black Forest Labs' unified multimodal frontier model — image, video, audio and action from a single set of weights. Its first generally available product, FLUX 3 Video, generates clips of up to 20 seconds with optional native audio from text, images or video.

4-in-1
Image, video, audio, action
HD–4K
Resolution tiers (BFL API)
20s
Native-audio video (same model)
2026
Announced July 23
What's new

One model for all visual intelligence

01

Unified multimodal core

Unlike FLUX.2, FLUX 3 learns jointly from images, video and audio in one architecture, then extends to action prediction. FLUX 3 Video draws on that shared understanding of scenes, motion and sound.

02

Up to 20 seconds per clip

FLUX 3 Video generates a single clip of up to 20 seconds, and through the BFL API it can also extend existing videos or transition between keyframes.

03

Native audio, same pass

Switch on native audio to generate sound together with the video in one pass — BFL doesn't charge extra for audio.

04

Self-Flow foundation

FLUX 3 is built on Self-Flow, BFL's flow-matching approach that aligns generation and understanding across modalities inside the same model.

Overview

Model at a glance

FLUX 3 is the third-generation flagship from Black Forest Labs (BFL), the Freiburg, Germany lab behind the FLUX family. Announced on 23 July 2026, it is a departure from its predecessors: rather than a standalone image model, FLUX 3 is a unified multimodal frontier model that jointly learns from images, video and audio within a single architecture and can be extended to predict actions. BFL frames it as a 'foundation layer for visual intelligence' — one network spanning image, video, audio and physical-AI action prediction, delivered across product lines — FLUX 3 Video and FLUX 3 Action today, with FLUX 3 Image and the open-weight FLUX 3 Dev on BFL's roadmap.

This page focuses on FLUX 3 Video, the first FLUX 3 product to reach general availability: it opened on the BFL API and select partner platforms on 4 August 2026. It covers text-to-video, image-to-video, video-to-video editing, video extension and keyframe transitions, generates up to 20 seconds per clip with optional native audio at no extra cost, and renders in HD, FHD, QHD and UHD-4K tiers. FLUX 3 Action, a 7B open-weight model for robotics, followed on 23 September 2026. FLUX 3 Image, for image generation and editing, has been announced but not yet released, with no public API as of September 2026.

Model
FLUX 3 Video
Developer
Black Forest Labs
Announced
23 July 2026
Type
Unified multimodal flow model
Foundation
Self-Flow (flow matching)
Video availability
Generally available (BFL API, Aug 4, 2026)
FLUX 3 Image
Announced; not yet released (no public API as of September 2026)
Open weights
FLUX 3 Dev, later in 2026
Predecessor
FLUX.2
Capabilities

Key features

Shared multimodal understanding

FLUX 3 is trained jointly across video, images and audio, so FLUX 3 Video draws on a shared model of scenes, motion and sound. BFL reports stronger complex-prompt handling in its preliminary mid-training evaluations.

Text, image and video inputs

Generate from a text prompt, animate a still image, or edit an existing clip with video-to-video. FLUX 3 Video also supports extending videos and keyframe transitions.

Optional native audio

FLUX 3 Video can add native audio to a clip in the same generation, so sound and picture come from one model instead of a separate audio step.

Self-Flow architecture

FLUX 3 is built on Self-Flow, BFL's approach that combines a flow-matching objective with self-supervised feature reconstruction to align generation and understanding inside a single model. Self-Flow was first introduced by BFL in March 2026.

Resolution tiers from HD to 4K

FLUX 3 Video renders in HD, FHD, QHD and UHD-4K tiers based on pixels per frame, with pricing on the BFL API that scales by tier.

Part of a coordinated model family

Video and action ship first as product lines over one backbone, with FLUX 3 Image for image generation and editing and the open-weight FLUX 3 Dev on BFL's roadmap; FLUX 3 Dev is planned for later in 2026.

Specs

Technical specifications

Model

NameFLUX 3 Video
DeveloperBlack Forest Labs (Freiburg, Germany)
Announced23 July 2026
Model typeUnified multimodal flow model
FoundationSelf-Flow (flow matching + self-supervised feature reconstruction)

Video capabilities

Modalities in one modelImage, video, audio, action
Generation modesText-to-video, image-to-video, video-to-video editing, extension, keyframe transitions
Max clip length20 seconds per generation
AudioOptional native audio (no extra charge)
Resolution tiersHD, FHD, QHD, UHD-4K (BFL API)

Availability

FLUX 3 VideoGenerally available via the BFL API and partners since Aug 4, 2026
FLUX 3 Action7B open weights released Sep 23, 2026 (FLUX Kommunity License)
FLUX 3 ImageAnnounced; not yet released (no public API as of September 2026)
FLUX 3 Dev (open weights)Planned later in 2026
SharkFotoText to Video & Image to Video: 720p/1080p, 5–20 s, native audio on/off
In practice

Use cases

Short-form social video

Generate clips of up to 20 seconds with native sound for Reels, Shorts and TikTok from a single prompt, in vertical or landscape formats.

Product and e-commerce videos

Animate a product photo into a lifestyle clip with image-to-video, using a start frame and an optional end frame to control how the shot opens and closes.

Concept and pre-visualization

Turn storyboards and key frames into moving previews to test pacing, camera moves and mood before committing to a full production.

Ads and promos with sound

Produce short ads and promos with native audio generated in the same pass, reducing the need for a separate sound step.

Extending and editing footage

Through the BFL API, FLUX 3 Video's video-to-video editing, extension and keyframe transitions can restyle or lengthen existing clips.

Generational leap

FLUX 3 vs FLUX.2

FLUX 3 is a fundamental shift from FLUX.2: where FLUX.2 is an image-only model family (FLUX.2 [dev] has 32B parameters), FLUX 3 is a unified multimodal backbone whose first release is FLUX 3 Video. FLUX 3 Image is still on BFL's roadmap, so no image-mode specs have been published yet.

FeatureFLUX.2FLUX 3NEW
ScopeImage generation onlyUnified image, video, audio and action
Architecture32B rectified-flow transformer + Mistral 3 VLMSelf-Flow multimodal flow model
ResolutionUp to 4MP (e.g. 2048x2048)Video: HD to UHD-4K tiers
Video & audioNot supportedUp to 20 s video with optional native audio
AvailabilityGenerally available (Pro/Flex/Dev)Video GA since Aug 2026; Image on the roadmap
Honest look

Current limitations

FLUX 3 Image not released yet

FLUX 3 Image, for image generation and editing, is on BFL's roadmap but had no public release or API as of September 2026; today FLUX 3 is available as FLUX 3 Video.

SharkFoto offers a subset

On SharkFoto, FLUX 3 renders 5–20 second clips at 720p or 1080p in Text to Video and Image to Video. The BFL API's QHD and 4K tiers, video-to-video editing, extension and keyframe transitions aren't offered here yet.

Open weights come later

The open-weight FLUX 3 Dev backbone for content creation is still pending; only the robotics-focused FLUX 3 Action weights are public so far.

Multimodal, not image-first

FLUX 3 is trained jointly across video, images and audio; teams that need a dedicated, production-proven image pipeline today may find FLUX.2 more predictable.

FAQ

Frequently asked questions

What is FLUX 3?
FLUX 3 is Black Forest Labs' third-generation flagship, announced on 23 July 2026. Unlike earlier FLUX releases, it is a unified multimodal frontier model trained jointly across video, images and audio, and it can be extended to predict physical actions. Its first generally available product is FLUX 3 Video, which generates clips of up to 20 seconds with optional native audio.
Is FLUX 3 an image or a video model?
Both are planned, but today FLUX 3 is available as a video model. FLUX 3 is a multimodal backbone delivered across product lines: FLUX 3 Video (generally available since August 2026), FLUX 3 Action (open weights, September 2026), and FLUX 3 Image and the open-weight FLUX 3 Dev, which are still on BFL's roadmap.
How is FLUX 3 different from FLUX.2?
FLUX.2 is an image-only model family — FLUX.2 [dev] is a 32B rectified-flow transformer that generates and edits images at up to 4 megapixels. FLUX 3 is a unified multimodal architecture built on BFL's Self-Flow approach, and its first release adds what FLUX.2 never had: video of up to 20 seconds with optional native audio.
Can FLUX 3 generate audio?
Yes. FLUX 3 Video can generate native audio together with the video in the same pass, and BFL doesn't charge extra for it. On SharkFoto you can switch native audio on or off for each generation.
Is FLUX 3 Image available yet?
Not yet. BFL's roadmap includes FLUX 3 Image for image generation and editing, but as of September 2026 it has no public release or API. FLUX 3 Video has been generally available via the BFL API and partners since 4 August 2026, FLUX 3 Action's 7B open weights were released on 23 September 2026, and FLUX 3 Dev open weights are planned for later in 2026.
What resolution and length does FLUX 3 Video produce?
Through the BFL API, FLUX 3 Video renders clips of up to 20 seconds in HD, FHD, QHD and UHD-4K tiers. On SharkFoto you can generate 5–20 second clips at 720p or 1080p. Its predecessor FLUX.2 is image-only and generates and edits at up to 4 megapixels.
Can I use FLUX 3 on SharkFoto?
Yes. FLUX 3 (FLUX 3 Video) is available on SharkFoto in Text to Video and Image to Video. Generate 5–20 second clips at 720p or 1080p in auto, 16:9, 9:16, 1:1, 4:3, 3:4 or 21:9, start from a first frame with an optional end frame, and switch native audio on or off.

Create with FLUX 3 on SharkFoto

Turn a prompt or a still image into a 5–20 second clip at up to 1080p with optional native audio — FLUX 3 Video is available now in SharkFoto's Text to Video and Image to Video tools.

Try FLUX 3 now