# Qwen-Image-3.0

> Source: https://aiwiki.ai/wiki/qwen_image_3
> Summary: Qwen-Image-3.0 is a text-to-image foundation model announced by Alibaba's Qwen team on July 21, 2026, as the third generation of the Qwen-Image series .
> Updated: 2026-08-07
> Fact-checked: 2026-08-07
> Categories: AI Models, Chinese AI, Generative AI, Image Generation
> License: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) - attribute to "AI Wiki (aiwiki.ai)"
> Cite as: AI Wiki. "Qwen-Image-3.0." aiwiki.ai, 7 Aug 2026. https://aiwiki.ai/wiki/qwen_image_3
> From AI Wiki (https://aiwiki.ai), the free encyclopedia of artificial intelligence. Reuse freely with attribution.

| Field | Value |
| --- | --- |
| Developer | Qwen team, [Alibaba](https://aiwiki.ai/wiki/alibaba) |
| Announced | July 21, 2026 |
| Type | [Text-to-image](https://aiwiki.ai/wiki/text_to_image) foundation model with image editing |
| Series | [Qwen-Image](https://aiwiki.ai/wiki/qwen_image), third generation |
| Claimed input length | Up to 4,500 tokens |
| Claimed language coverage | Native text rendering in 12 languages |
| Parameters | Not disclosed |
| Weights | Not released (as of August 7, 2026) |
| License | None published |
| Access | Qwen Chat; API (invitational testing at launch) |

**Qwen-Image-3.0** is a [text-to-image](https://aiwiki.ai/wiki/text_to_image) foundation model announced by [Alibaba](https://aiwiki.ai/wiki/alibaba)'s [Qwen](https://aiwiki.ai/wiki/qwen) team on July 21, 2026, as the third generation of the [Qwen-Image](https://aiwiki.ai/wiki/qwen_image) series [1][2][15]. The Qwen team organized the release around the single keyword "Real" (实), broken into three claimed strengths: "Rich Content" (prompts up to 4,500 tokens that yield dense multi-panel layouts such as newspapers, storyboards, and exam papers), "Authentic Details" (legible text down to 10 pixels, plus fine textures like pores and hair strands), and "Deep Knowledge" (native text rendering in 12 languages, simulation of web, game, and livestream interfaces, and use of world knowledge) [1][2]. Unlike the original 2025 Qwen-Image, which shipped its weights under the [Apache 2.0](https://aiwiki.ai/wiki/apache_license) license with a same-day technical report, Qwen-Image-3.0 launched as a hosted model only: no weights, license, parameter count, benchmark table, or technical report accompanied the announcement [3][4].

## Background: the Qwen-Image series

The Qwen-Image line began in August 2025 as an open-weight [image generation](https://aiwiki.ai/wiki/image_generation) model. The original release was a roughly 20-billion-parameter multimodal [diffusion transformer](https://aiwiki.ai/wiki/diffusion_transformer) with an arXiv technical report, distributed on [Hugging Face](https://aiwiki.ai/wiki/hugging_face) and ModelScope under Apache 2.0, and it became known above all for rendering accurate Chinese and English text inside images [5][6][7]. A string of open-weight follow-ups extended it through 2025: the Qwen-Image-Edit editing models (August and September 2025), the Qwen-Image-Layered decomposition model and Qwen-Image-Edit-2511 (December 2025), and an improved base model, Qwen-Image-2512 (December 2025), all under Apache 2.0 [7][8][9].

The series changed course in February 2026 with Qwen-Image-2.0, a "next-generation foundational image generation model" that unified generation and editing in one system, accepted instructions of about 1,000 tokens, supported native 2K output, and used what the team described as a lighter architecture with faster inference than its predecessor [10]. Qwen-Image-2.0 was offered through Qwen Chat and Alibaba's APIs rather than as open weights, and the team evaluated it by blind testing on Alibaba's internal AI Arena leaderboard [10]. Qwen-Image-3.0 follows the same closed pattern [3][4].

The Qwen team summarizes the generations by keyword: 1.0 stood for "Precision," 2.0 for "Precision, Variety, Completeness, Beauty, and Authenticity," and 3.0 for "Real," which the team glosses as moving image generation "from 'good-looking' to 'useful'" so that it works as "a truly deployable productivity tool" [1].

| Release | Date | Distribution | License |
| --- | --- | --- | --- |
| Qwen-Image (1.0) | August 4, 2025 | Open weights (Hugging Face, ModelScope) | Apache 2.0 [5][7] |
| Qwen-Image-Edit | August 18, 2025 | Open weights | Apache 2.0 [8] |
| Qwen-Image-Edit-2509 | September 2025 | Open weights | Apache 2.0 [8] |
| Qwen-Image-Layered / Edit-2511 | December 2025 | Open weights | Apache 2.0 [8] |
| Qwen-Image-2512 | December 2025 | Open weights | Apache 2.0 [9] |
| Qwen-Image-2.0 | February 2026 | Hosted only (Qwen Chat, API) | None published [10] |
| Qwen-Image-3.0 | July 21, 2026 | Hosted only (Qwen Chat, API) | None published [1][3] |

## Claimed capabilities

All capability claims below come from the Qwen team's launch post and its Chinese-language announcement; the company published no benchmark results or technical report against which to check them, and press coverage noted that the demonstrations are outputs Alibaba itself selected [1][2][3].

### Rich content

Qwen-Image-3.0 raises the maximum instruction length to 4,500 tokens, from roughly 1,000 tokens in Qwen-Image-2.0, which the team says lets the model render information-dense layouts such as newspaper pages, comic storyboards, and exam papers in a single pass [1][3]. The flagship demonstration is a 3x3 grid of nine unrelated infographics (covering topics from the Sylow theorems of group theory to a parasitology explainer and a bank internal-control chart) generated as one image from a 3,700-token prompt [1][2]. A second demonstration nests interfaces inside one another: a VS Code window containing a Qwen Chat screen, which contains a messaging-app conversation, which contains a coffee poster, with each layer keeping its own interface style [1][2].

### Authentic details

The team claims precise rendering of text as small as 10 pixels, shown through a whale-shark infographic, a full page of an algebraic-geometry paper with LaTeX formula derivations, and a newspaper page with simulated print texture [1]. Editing demonstrations include overlaying handwritten-style red annotations on a book page and restoring a damaged ink-wash painting while matching the original brushwork [1]. Portrait examples emphasize pores, hair strands, and skin texture that the team describes as approaching photographic realism [1][2].

### Deep knowledge

The model is said to render text natively in 12 languages (demonstrations show Japanese, Korean, and Spanish), across more than 20 fonts and over 100 artistic styles, and to imitate mainstream web, game, and livestream interfaces [1][2]. The team also states the hosted model can retrieve current information from the internet, demonstrated by generating a weather-forecast graphic for Hangzhou for a specific date, and can draw on knowledge of real figures, shown by placing the painters Qi Baishi and Vincent van Gogh in a livestream-room scene [1].

## Availability

At launch, the official announcement said API access was open for invitational testing on Alibaba Cloud's Bailian platform and the Qwen AI platform, with the Qwen Studio desktop application and the Qwen mobile app to add free access later [2][16]. The English blog post points readers to the text-to-image mode of Qwen Chat, and Decrypt reported the model live in chat.qwen.ai with API pricing not yet announced [1][4]. By early August 2026, a model named "qwen-image-3.0-pro" was listed on Alibaba's Qwen cloud console and in public arena testing; Alibaba has not published a breakdown of how the "Pro" API variant relates to the model described in the launch post [11].

## Benchmarks and leaderboard results

The launch post contains no benchmark table, and no Qwen-Image-3.0 technical report had been published as of August 7, 2026, so the only quantitative evidence comes from third-party arenas [1][3].

On [LMArena](https://aiwiki.ai/wiki/lmarena)'s Text-to-Image leaderboard, as of the board's August 4, 2026 update, "qwen-image-3.0-pro" ranked 5th of 75 models with an Arena score of about 1263 from 2,801 votes, behind gpt-image-2 (medium) at about 1380, reve-2.1, muse-image, and reve-2.0, and just ahead of gemini-3.1-flash-image ([Nano Banana](https://aiwiki.ai/wiki/nano_banana) 2) and [Seedream](https://aiwiki.ai/wiki/seedream) 5.0 Pro [11]. The entry's vote count was still low, giving it a wide confidence range (rank 4 to 9), and the prior flagship, qwen-image-2.0-pro-2026-06-22, ranked 15th at about 1191 [11].

For context, the Qwen team's own creator-centric benchmark, Qwen-Image-Bench, published in May 2026 with an accompanying paper, placed the previous flagship Qwen Image 2.0 Pro 5th of 18 models (overall score 57.84), behind [GPT Image](https://aiwiki.ai/wiki/gpt_image) 2 (64.69), Nano Banana 2.0, GPT Image 1.5, and Nano Banana Pro, and ahead of Seedream 5.0 and [FLUX](https://aiwiki.ai/wiki/flux) 2 Max [12][13]. Qwen-Image-3.0 postdates that evaluation, and no Qwen-Image-Bench score for it had been published as of August 7, 2026 [12].

## Openness and reception

Qwen-Image-3.0's closed release drew attention because of the series' open-weight history. Unite.AI wrote that "the post carries no benchmark table, no parameter count, no license, and no downloadable weights, and no technical report describing how the model was trained or tested," calling it "a departure from how the series shipped before" [3]. As of August 7, 2026, the Qwen organizations on Hugging Face and ModelScope listed no Qwen-Image-3.0 repository, leaving Qwen-Image-2512 from December 2025 as the newest open-weight base model in the line [9]. The Decoder judged it "unlikely that the model weights will ship under an open license, as they did for the original Qwen-Image" [14].

Coverage of the capabilities themselves was cautiously positive. The Decoder highlighted the single-pass infographic grids and 10-pixel text but noted that "whether AI-generated academic papers and newspaper pages are useful as static images remains an open question" [14]. Decrypt framed the release as Alibaba steering image generation toward practical document and interface work rather than aesthetics [4]. Unite.AI observed that without weights or an evaluation set, "the only evidence a developer can act on is Alibaba's own reel of outputs" [3].

## See also

- [Qwen-Image](https://aiwiki.ai/wiki/qwen_image)
- [Qwen](https://aiwiki.ai/wiki/qwen)
- [Text-to-image](https://aiwiki.ai/wiki/text_to_image)
- [Image generation](https://aiwiki.ai/wiki/image_generation)
- [Seedream](https://aiwiki.ai/wiki/seedream)
- [GPT Image](https://aiwiki.ai/wiki/gpt_image)
- [Nano Banana](https://aiwiki.ai/wiki/nano_banana)
- [Diffusion models](https://aiwiki.ai/wiki/diffusion_models)

## References

1. Qwen Team. "Qwen-Image-3.0: Rich Content, Authentic Details, Deep Knowledge." Qwen blog, July 21, 2026. https://qwen.ai/blog?id=qwen-image-3.0
2. IT之家. "阿里千问发布 Qwen-Image-3.0 图像生成基础模型：落字成画，字字如印." July 21, 2026. https://www.ithome.com/0/979/530.htm
3. Unite.AI. "Alibaba Launches Qwen-Image-3.0 Without Benchmarks or Weights." July 21, 2026. https://www.unite.ai/alibaba-launches-qwen-image-3-0-without-benchmarks-or-weights/
4. Decrypt. "Alibaba's New Qwen Image 3 AI Wants to Be Useful, Not Just Pretty." July 22, 2026. https://decrypt.co/374084/alibaba-qwen-image-3-ai-useful-not-just-pretty
5. Qwen Team. "Qwen-Image: Crafting with Native Text Rendering." Qwen blog, August 4, 2025. https://qwenlm.github.io/blog/qwen-image/
6. Qwen Team. "Qwen-Image Technical Report." arXiv:2508.02324, August 2025. https://arxiv.org/abs/2508.02324
7. "Qwen/Qwen-Image." Hugging Face model card. https://huggingface.co/Qwen/Qwen-Image
8. "Qwen/Qwen-Image-Edit," "Qwen/Qwen-Image-Edit-2509," "Qwen/Qwen-Image-Edit-2511," "Qwen/Qwen-Image-Layered." Hugging Face model cards (Apache 2.0). https://huggingface.co/Qwen/Qwen-Image-Edit
9. "Qwen/Qwen-Image-2512." Hugging Face model card (Apache 2.0, created December 30, 2025). https://huggingface.co/Qwen/Qwen-Image-2512
10. Qwen Team. "Qwen-Image-2.0: Professional infographics, exquisite photorealism." Qwen blog, February 2026. https://qwen.ai/blog?id=qwen-image-2.0
11. "Text-to-Image Leaderboard." LMArena, leaderboard update of August 4, 2026 (retrieved August 7, 2026). https://arena.ai/leaderboard/text-to-image
12. "Qwen/Qwen-Image-Bench." Hugging Face dataset (leaderboard README, May 2026). https://huggingface.co/datasets/Qwen/Qwen-Image-Bench
13. Qwen Team. "Qwen-Image-Bench: From Generation to Creation in Text-to-Image Evaluation." arXiv:2605.28091, May 2026. https://arxiv.org/abs/2605.28091
14. The Decoder. "Alibaba's Qwen-Image-3.0 renders full infographic grids and readable ten-pixel text in a single pass." July 21, 2026. https://the-decoder.com/alibabas-qwen-image-3-0-renders-full-infographic-grids-and-readable-ten-pixel-text-in-a-single-pass/
15. Qwen (@Alibaba_Qwen). Announcement post on X, July 22, 2026. https://x.com/Alibaba_Qwen/status/2079906336381509659
16. 新浪财经. "阿里发布Qwen-Image-3.0图像模型，支持超长指令复杂图文生成." July 21, 2026. https://finance.sina.com.cn/tech/shenji/2026-07-21/doc-iniipxxk4931252.shtml

