Alibaba · Real over pretty: images that work as documents.

Qwen-Image 3.0

Text-Heavy Images, Rendered Right

Alibaba's Qwen-Image 3.0 is a text-to-image model built for dense, information-rich layouts — newspapers, infographics, and mockups with readable type, generated in a single pass.

4,500
Max prompt tokens
12
Languages
20+
Fonts
~10px
Legible text size
What's new

Text-Heavy Images, Rendered Right

01

4,500-Token Prompts

An ultra-long instruction window — roughly 4.5x the ~1,000-token limit of Qwen-Image 2.0 — lets you specify dense, multi-element layouts in a single prompt.

02

Readable Text at ~10px

Alibaba shows legible characters as small as ten pixels, plus LaTeX math with subscripts, superscripts, and complex notation.

03

Full Infographic Grids

Renders multi-panel layouts — demonstrated as 3x3 grids holding nine separate infographics — correctly composed in one generation.

04

Multilingual & Multi-Font

Native support for 12 languages and 20+ fonts, aimed at practical documents rather than purely aesthetic imagery.

Overview

Model at a glance

Qwen-Image 3.0 is the third-generation text-to-image model from Alibaba's Qwen (Tongyi) team, released on July 21, 2026. Where most image generators chase aesthetics, Qwen-Image 3.0 pitches itself around usefulness: its stated goal is 'Real,' targeting practical work such as newspaper layouts, storyboards, exam sheets, and multi-panel infographics that need correct structure and legible text to function as working documents. The headline capability is an ultra-long prompt window of up to 4,500 tokens — roughly 4.5x the limit of Qwen-Image 2.0 — which lets a single prompt describe dense, multi-element compositions the model composes in one pass.

The model supports 12 languages and more than 20 fonts, renders text legible down to about ten pixels, and handles LaTeX mathematical notation with subscripts, superscripts, and complex expressions. Alibaba's launch gallery demonstrates 3x3 grids of nine distinct infographics, nested interface mockups, photorealistic portraits with visible skin texture and individual hair strands, and image editing and restoration. Notably, Qwen-Image 3.0 broke from the open-source pattern of Qwen-Image 1.0 and 2.0: it shipped with no open weights, no benchmark scores, no technical report, and no model card — available only through Alibaba's hosted apps and API. As a result, its quality claims rest on the company's own curated examples rather than independent evaluation.

Developer
Alibaba — Qwen (Tongyi) team
Model type
Text-to-image generation
Release date
July 21, 2026
Prompt window
Up to 4,500 tokens
Languages
12 supported
Fonts
20+
Weights
Not released (closed / hosted)
Access
Qwen Chat, Qwen Studio, Alibaba API
Capabilities

Key features

Ultra-long prompt window

Accepts prompts of up to 4,500 tokens — a large jump from the roughly 1,000-token limit of Qwen-Image 2.0 — so a single instruction can specify a dense, multi-region layout with per-element content, styling, and copy.

Small, legible text rendering

Alibaba demonstrates readable characters as small as about ten pixels, a threshold most image models miss, making the model suited to captions, labels, tables, and fine print inside a composition.

Multi-panel infographic layouts

Generates structured grids — shown as 3x3 arrangements of nine separate infographics — with each panel correctly composed, rather than a single centered subject on a background.

Multilingual and multi-font typography

Native handling of 12 languages, including Japanese, Korean, and Spanish, across 20+ fonts, targeting documents that mix scripts and typefaces.

LaTeX and mathematical notation

Renders LaTeX-style formulas with subscripts, superscripts, and complex mathematical expressions, useful for exam sheets, worksheets, and technical diagrams.

Photorealistic detail

Alibaba's samples show fine textures — visible skin detail, individual strands of hair, and paper grain — reproduced at near-photographic quality alongside the text-layout focus.

Editing and restoration

Beyond generation, the model is shown editing images and restoring damaged material, including repairing traditional ink paintings, and integrating live data into generated graphics.

Specs

Technical specifications

Model

DeveloperAlibaba — Qwen (Tongyi) team
GenerationThird generation (Qwen-Image line)
ModalityText-to-image
Release dateJuly 21, 2026
Parameter countNot disclosed

Prompting & Text

Max prompt lengthUp to 4,500 tokens
Languages12
Fonts20+
Smallest legible text~10 pixels
Math notationLaTeX (sub/superscripts, complex expressions)

Access & Availability

Open weightsNo (closed / hosted)
Technical reportNot published
BenchmarksNone released at launch
Access channelsQwen Chat, Qwen Studio, Alibaba API (invite-only at launch)
Public pricingNot published at launch
In practice

Use cases

Infographics and data graphics

Produce multi-panel infographics, charts, and weather-style data cards where each region needs correct labels, numbers, and structure in one pass.

Document and poster layouts

Generate newspaper-style pages, posters, and flyers with dense, legible copy across multiple fonts and languages.

Educational materials

Create exam sheets, worksheets, and study aids that require readable small text and correctly rendered LaTeX math.

Interface and product mockups

Render nested UI mockups — editor windows, chat apps, and app screens — with plausible on-screen text and layout for concepting.

Storyboards and photorealistic portraits

Build storyboard frames or detailed portraits that combine fine texture with any accompanying captions or annotations.

Generational leap

FeatureQwen-Image 2.0Qwen-Image 3.0NEW
~1,000 tokensUp to 4,500 tokens
Yes (Apache 2.0)No (closed / hosted)
Same-day reportNot published
Published (Qwen-Image-Bench)None released
General image qualityDense text & practical layouts
Honest look

Current limitations

No independent benchmarks

Qwen-Image 3.0 launched with no benchmark scores, technical report, or model card. All quality claims rest on Alibaba's own curated example outputs rather than third-party evaluation.

Closed and hosted only

Unlike Qwen-Image 1.0 and 2.0, version 3.0 ships no open weights and cannot be self-hosted. Access is limited to Alibaba's Qwen Chat, Qwen Studio, and API, invite-only at launch.

No published pricing or parameter count

At launch Alibaba disclosed neither pricing nor a parameter count, making cost planning and capability comparison difficult.

Claims unverified in the wild

Headline figures — the 4,500-token window, ~10px legible text, and nine-panel grids — reflect demonstrated examples; real-world reliability across arbitrary prompts has not been independently confirmed.

FAQ

Frequently asked questions

What is Qwen-Image 3.0?
It is the third-generation text-to-image model from Alibaba's Qwen (Tongyi) team, released July 21, 2026. It is designed for dense, text-heavy, practical layouts — infographics, documents, and mockups — rather than purely aesthetic imagery.
How is Qwen-Image 3.0 different from Qwen-Image 2.0?
It expands the prompt window from roughly 1,000 to 4,500 tokens and centers on legible dense text and structured layouts. Critically, unlike 2.0 (which shipped open weights under Apache 2.0 with a same-day technical report), 3.0 is closed — no weights, no benchmarks, and no report.
Is Qwen-Image 3.0 open source?
No. Qwen-Image 3.0 is a closed, hosted model. Alibaba did not release weights, a technical report, or a model card, and it cannot be self-hosted.
How many languages and fonts does it support?
Alibaba states native support for 12 languages and more than 20 fonts, alongside legible text rendering down to about ten pixels and LaTeX mathematical notation.
How can I access Qwen-Image 3.0?
Through Alibaba's hosted channels — Qwen Chat and Qwen Studio apps and the Alibaba API. Access was invite-only at launch, with no public pricing announced.
Does Qwen-Image 3.0 have published benchmark scores?
No. It launched without benchmark scores or a technical report, so its quality claims are based on Alibaba's own curated example outputs rather than independent evaluation.
Can I use Qwen-Image 3.0 on SharkFoto?
Not yet — Qwen-Image 3.0 is coming soon to SharkFoto as integration is a later step. In the meantime you can use SharkFoto's currently available AI image tools for generation, editing, and enhancement, and check back for Qwen-Image 3.0 support.

Qwen-Image 3.0 is coming to SharkFoto

We're preparing to bring Alibaba's text-heavy Qwen-Image 3.0 to SharkFoto. It isn't live yet — explore SharkFoto's available AI image tools now, and check back soon for Qwen-Image 3.0.

Try Qwen-Image 3.0 now