MAI-Image 2.5 is Microsoft AI's first first-party text-to-image and editing model, pairing breakthrough text rendering with surgical, identity-preserving edits.
A +107 Arena gain over MAI-Image 2 makes legible signage, labels, and multilingual captions a core strength rather than an afterthought.
Launched at No. 2 on Arena's image-editing leaderboard, ahead of Nano Banana 2.1, with localized edits that leave the rest of the frame untouched.
Faces, hair, clothing, and full-body identity hold across stylization, pose, and layout changes for consistent characters.
Choose fast-and-cheap Flash, the balanced 2.5 standard, or the highest-fidelity 2.5-Pro tier for portrait-grade detail.
MAI-Image 2.5 is Microsoft AI's first fully in-house text-to-image and image-editing model, marking the company's move to power its own creative products rather than lean on third-party generators. Launched on June 2, 2026, it debuted at No. 3 on Arena's text-to-image leaderboard and No. 2 for image editing, ahead of Nano Banana 2.1. Against Microsoft's own prior MAI-Image 2, it posts an overall +75-point Arena gain, led by a +107 jump in text rendering and +90 in the cartoon, anime, and fantasy category.
Beyond raw quality, the model is built for controllable, production-grade workflows: localized edits that change only what you ask, face and identity consistency across stylization and pose, and structured document and diagram generation that yields presentation-ready visuals. It ships as a family — a low-cost Flash tier, the balanced 2.5 standard, and the higher-fidelity 2.5-Pro tier previewed in July 2026 that Microsoft bills as its highest-fidelity image model to date. The line already powers Bing Image Creator end-to-end and drives image features in PowerPoint and OneDrive, and it is available to developers through Azure AI Foundry, the MAI Playground, and OpenRouter.
Microsoft calls its text accuracy a real breakthrough, with roughly a +107 Arena ELO gain over MAI-Image 2. The model handles signage, product labels, posters, and multilingual captions that stay sharp and readable.
Object removal and replacement, attribute changes, inpainting, blur removal, and artifact cleanup are applied precisely, preserving the surrounding composition instead of regenerating the whole image.
Recognizable faces, plus hair, clothing, and full-body identity, are maintained across stylization, pose, and layout changes so characters stay coherent across a series of images.
Iterate on an image with plain-language instructions rather than masks and tools, making creative revision faster and more intuitive for non-experts.
Structured, presentation-ready visuals and diagrams are a first-class capability, which is why the line drives image generation inside PowerPoint.
The model composes coherent scenes from vague prompts, reasoning about structure, lighting, spatial relationships, and the reflection and weight of materials for believable results.
A Flash tier prioritizes speed and low cost, the 2.5 standard balances both, and 2.5-Pro maximizes fidelity for high-definition portraits and detailed scenes.
| Developer | Microsoft AI (MAI) |
|---|---|
| Type | Text-to-image + image editing |
| Tiers | MAI-Image 2.5-Flash, 2.5, 2.5-Pro |
| First release | June 2, 2026 |
| Pro tier preview | July 23, 2026 |
| Generation | Text-to-image, image-to-image |
|---|---|
| Editing | Localized edits, inpainting, object replace, cleanup |
| Text rendering | High accuracy, multilingual |
| Consistency | Face + full-body identity across pose/style |
| Image editing | No. 2 at launch |
|---|---|
| Text-to-image | No. 3 at launch |
| Gain vs MAI-Image 2 | +75 overall Arena |
| Text-rendering gain | +107 vs MAI-Image 2 |
| Platforms | Azure AI Foundry, MAI Playground, OpenRouter |
|---|---|
| Text input | $5 / 1M tokens |
| Image input | $8 / 1M tokens |
| Image output | $47 / 1M tokens (standard) |
Generate clean product shots, posters, and ad creatives with sharp, legible on-image text and correct branding-style layouts.
Produce presentation-ready diagrams, illustrations, and scene compositions directly inside slide and document workflows.
Apply localized edits such as object removal, replacement, and blur cleanup while keeping the rest of the photo untouched.
Keep a person or character recognizable across stylizations, poses, and layouts for storyboards, comics, and campaign sets.
Lean on the strong cartoon, anime, and fantasy gains to explore stylized worlds and characters from loose prompts.
| Feature | MAI-Image 2 | MAI-Image 2.5NEW |
|---|---|---|
| Baseline | +75 points | |
| Baseline | +107 points | |
| Baseline | +90 points | |
| Limited control | Surgical, identity-preserving | |
| Lower tier | No. 2 at launch |
Microsoft has not publicly detailed exact maximum output resolutions, aspect-ratio limits, or a full supported-language list, so those parameters should be confirmed in Azure AI Foundry documentation before production use.
MAI-Image 2.5-Pro entered public preview in July 2026 and carries a notably higher image-output price than the standard and Flash tiers, so quality gains come with cost and stability trade-offs.
The model runs through Microsoft and partner platforms rather than as an open-weight download, so usage is tied to those providers' terms and quotas.
Like all image models, it can misrender fine details, small text, hands, or complex spatial relationships, and outputs should be reviewed before publishing.
Microsoft's first-party image model — with breakthrough text rendering and surgical editing — is on its way to SharkFoto. Explore SharkFoto's available AI image tools today, and check back here to be first to try MAI-Image 2.5 when it launches.
Try MAI-Image 2.5 now