FLUX 3 is Black Forest Labs' unified multimodal frontier model — image, video, audio and action from a single set of weights. Its Image mode targets photoreal detail, any style, and typographically accurate multilingual text.
Unlike FLUX.2, FLUX 3 learns jointly from images, video and audio in one architecture, then extends to action prediction. Image generation draws on that shared understanding rather than a separate image-only network.
FLUX 3 renders in-image text — signs, labels, posters, UI mockups — with far fewer garbled characters than prior FLUX and Stable Diffusion releases, including dense copy and non-Latin scripts.
BFL positions FLUX 3 Image to close most of the remaining photorealism gap with Midjourney on skin texture, lighting and multi-subject scenes, while spanning illustration, product renders and fine art.
FLUX 3 is built on Self-Flow, BFL's flow-matching approach that aligns generation and understanding across modalities inside the same model.
FLUX 3 is the third-generation flagship from Black Forest Labs (BFL), the Freiburg, Germany lab behind the FLUX family. Announced on 23 July 2026, it is a departure from its predecessors: rather than a standalone image model, FLUX 3 is a unified multimodal frontier model that jointly learns from images, video and audio within a single architecture and can be extended to predict actions. BFL frames it as a 'foundation layer for visual intelligence' — one network spanning image, video, audio and physical-AI action prediction, delivered across product lines including FLUX 3 Image, FLUX 3 Video, FLUX 3 Action and the upcoming open-weight FLUX 3 Dev.
This page focuses on FLUX 3's image capabilities. Because the model shares one backbone across modalities, image generation benefits from the same multimodal understanding used for video and audio, which BFL credits for stronger complex-prompt handling and markedly better in-image text rendering than FLUX.2. Note that FLUX 3 rolled out in phases: FLUX 3 Video and FLUX 3 Action opened in gated early access at launch, while FLUX 3 Image was announced as arriving 'in the coming weeks' and is only entering early access around early August 2026. As a result, some image-specific numbers (maximum resolution, parameter count) have not yet been officially published, and current image-quality claims come from BFL's own mid-training evaluations.
FLUX 3 generates images from the same weights that model video and audio, so the image mode inherits a richer world model. BFL reports this yields better handling of complex, compositional prompts than FLUX.2's image-only design.
One of FLUX 3's headline improvements is typography. It renders signs, labels, posters and UI copy with far fewer garbled characters, and extends to dense text blocks and non-Latin scripts — a long-standing weak point for diffusion image models.
BFL positions FLUX 3 Image to hold up on skin texture, lighting and multi-subject scenes at a level closer to Midjourney, while remaining a generalist that spans photography, illustration, product renders and fine art.
FLUX 3 is built on Self-Flow, BFL's approach that combines a flow-matching objective with self-supervised feature reconstruction to align generation and understanding inside a single model. Self-Flow was first introduced by BFL in March 2026.
FLUX 3 Image is described as supporting a wide output range at flexible aspect ratios and resolutions, from square social crops to wide landscape formats, though exact maximum-resolution figures are not yet officially published for the image mode.
Image, video, audio and action ship as product lines over one backbone, so a prompt style and understanding carry across modalities. An open-weight FLUX 3 Dev backbone is planned for later in 2026 for self-hosting and research.
| Name | FLUX 3 (Image mode) |
|---|---|
| Developer | Black Forest Labs (Freiburg, Germany) |
| Announced | 23 July 2026 |
| Model type | Unified multimodal flow model |
| Foundation | Self-Flow (flow matching + self-supervised feature reconstruction) |
| Modalities in one model | Image, video, audio, action |
|---|---|
| Text rendering | Multilingual, including dense text and non-Latin scripts |
| Style range | Photography, illustration, product renders, fine art |
| Aspect ratios | Flexible (square to wide landscape) |
| Max image resolution | Not yet officially published for FLUX 3 Image |
| FLUX 3 Video / Action | Gated early access at launch |
|---|---|
| FLUX 3 Image | Entering early access (rolling out from ~Aug 2026) |
| FLUX 3 Dev (open weights) | Planned later in 2026 |
| SharkFoto | Coming soon |
FLUX 3's stronger text rendering makes it well suited to layouts that need legible headlines, labels and body copy baked directly into the image, reducing manual retyping in an editor.
Photoreal lighting and material handling support clean product shots and lifestyle scenes, with flexible aspect ratios for storefront, ad and social placements.
The generalist style range spans illustration, fine art and stylised looks, giving artists and studios a single model for mood boards, key art and iteration.
Accurate rendering of non-Latin scripts helps teams create localized signage, packaging mockups and campaign assets across markets without swapping models.
Cleaner rendering of UI labels and dense text makes FLUX 3 useful for early-stage app and web mockups where readable interface copy matters.
FLUX 3 is a fundamental shift from FLUX.2: where FLUX.2 was a 12B-parameter image-only model, FLUX 3 is a unified multimodal backbone. Note that FLUX 3 Image is still pre-general-availability, so several of its image specs are not yet officially confirmed.
| Feature | FLUX.2 | FLUX 3NEW |
|---|---|---|
| Scope | Image generation only | Unified image, video, audio and action |
| Architecture | 12B-param hybrid diffusion transformer | Self-Flow multimodal flow model |
| Native resolution | Up to 4K (3840x2160), 4MP native | Not yet officially published for Image mode |
| In-image text | Strong for a diffusion model | Improved multilingual accuracy, incl. non-Latin scripts |
| Availability | Generally available (Pro/Flex/Dev) | Phased early access; Image rolling out from ~Aug 2026 |
At the 23 July 2026 launch, only FLUX 3 Video and FLUX 3 Action opened in early access. FLUX 3 Image is entering gated early access around August 2026, so broad, self-serve access may still be limited.
BFL has not yet released official image-specific numbers such as maximum resolution or parameter count for FLUX 3 Image. Current image-quality claims are based partly on BFL's own mid-training evaluations rather than fully independent testing.
The open-weight FLUX 3 Dev backbone is planned for later in 2026. Until then, self-hosting and offline fine-tuning of FLUX 3 are not available.
FLUX 3 is engineered as a unified backbone where video prediction received the vast majority of training compute. Teams that want a dedicated, battle-tested image pipeline today may find FLUX.2 more predictable until FLUX 3 Image matures.
We're adding FLUX 3 as platform access opens — explore SharkFoto's available AI tools today and check back for availability.
Try FLUX 3 now