# Ming-Image-0.1-Design

> Source: https://aiwiki.ai/wiki/ming_image_0_1_design
> Updated: 2026-09-25
> Fact-checked: 2026-09-25
> Categories: Chinese AI, Diffusion Models, Image Generation, Open Source AI
> License: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) - attribute to "AI Wiki (aiwiki.ai)"
> Cite as: AI Wiki. "Ming-Image-0.1-Design." aiwiki.ai, 25 Sept 2026. https://aiwiki.ai/wiki/ming_image_0_1_design
> From AI Wiki (https://aiwiki.ai), the free encyclopedia of artificial intelligence. Reuse freely with attribution.

**Ming-Image-0.1-Design** is a family of open-weight image models for graphic design released by [inclusionAI](https://aiwiki.ai/wiki/inclusionai), the open-source AI initiative founded by [Ant Group](https://aiwiki.ai/wiki/ant_group), and announced under Ant Group's Ant Ling brand on 22 September 2026.[1][2] The family has two models, each described by the developer as 6 billion parameters: Ming-Image-0.1-Design, a [text-to-image](https://aiwiki.ai/wiki/text-to-image_models) model for user interfaces, dashboards, infographics, posters and other text-heavy layouts that can also output images with a transparent alpha channel, and Ming-Image-0.1-Design-Layer, which splits a flattened design image into 2 to 9 separately editable RGBA layers.[1][3][4] Ant released two companion Agent Skills with the models, the Ling UI Design Skill and the Image-to-Editable-PPT Skill.[1][5] The weights are published on [Hugging Face](https://aiwiki.ai/wiki/hugging_face) and [ModelScope](https://aiwiki.ai/wiki/modelscope) under the MIT License.[3][4][6]

At launch Ant Ling said Ming-Image-0.1-Design ranked first among open-weight models on [Artificial Analysis](https://aiwiki.ai/wiki/artificial_analysis)'s UI/UX Design leaderboard.[1] When checked on 25 September 2026, the model was still first in that category's open-weights view, with an Elo score of 1,084, and 16th among all models in the category.[7][8]

## Background

inclusionAI's website says it was founded by and is backed by Ant Group, and lists "Open by Default" among its principles.[9] Ant Ling (@AntLingAGI on X) describes itself as the "MoE model series with foundation (Ling), reasoning (Ring) and any-to-any (Ming) from Ant Group's AGI initiative", and links to inclusionAI's account.[1] The Chinese launch post was published on WeChat by the Bailing large-model account (百灵大模型).[2] The model repositories are hosted under the inclusionAI organization on Hugging Face and ModelScope.[3][4][6]

Ming is inclusionAI's multimodal line. Its June 2025 paper, "Ming-Omni: A Unified Multimodal Model for Perception and Generation" (arXiv 2506.09344), described a model that processes images, text, audio and video with Ling, an MoE ([mixture of experts](https://aiwiki.ai/wiki/mixture_of_experts)) architecture with modality-specific routers, and that can generate speech and images.[10] Ming-Image-0.1-Design is narrower: it targets design images rather than general multimodal chat.[2][3]

In its launch post the team framed the release around two Ant Group products. In flash apps (闪应用) in the Lingguang (灵光) app, users had to wait until an app finished generating before they could see its page design. In Huamei (画眉), which the post describes as Ant Group's internal intelligent-design application for designers and business teams, designers usually needed more than 5 minutes of layer-splitting work on a generated image before they could keep editing it.[2] The post says the goal was to move AI design from generating a picture toward output that can be previewed, edited and delivered.[2]

## Models

| Model | Size (developer figure) | Task | Hugging Face pipeline tag | Licence | Repository created |
| --- | --- | --- | --- | --- | --- |
| Ming-Image-0.1-Design | 6B | Text-to-image design generation, including transparent RGBA output | text-to-image | MIT | 17 September 2026 |
| Ming-Image-0.1-Design-Layer | 6B | Decomposes a flattened design image into 2-9 RGBA layers from an image plus a layer plan | image-text-to-image | MIT | 17 September 2026 |

Sources: model cards, Hugging Face API records and the launch thread.[1][3][4][11]

Both repositories contain an MIT License file with the notice "Copyright (c) 2026 inclusionAI".[11] The Hugging Face repositories were created on 17 September 2026; the public announcement followed on 22 September (UTC), and OpenRouter lists the Design model's release date as 22 September 2026.[1][11][12] The Chinese blog post is dated 23 September 2026 Beijing time.[2]

### Architecture and parameter counts

The model cards do not describe the architecture beyond the 6B figure, and none of the launch materials links a technical report. The checkpoint files show the structure. Each repository contains five components: `mllm/` (a multimodal language model), `mlp/`, `connector/`, `transformer/` (the diffusion model) and `vae/`, plus a flow-matching scheduler (`FlowMatchEulerDiscreteScheduler`).[11]

| Component | What the configuration files show | Parameters |
| --- | --- | --- |
| `mllm/` | `BailingMM2NativeForConditionalGeneration`: a 20-layer `BailingMoeV2` language model with 256 experts, 8 active per token and one shared expert, plus a `Qwen2_5_VisionTransformer` vision encoder | 17,001,125,376 (index metadata) |
| `connector/` | A 28-layer model using the `Qwen2ForCausalLM` architecture (hidden size 1,536) | 1,543,714,304 (index metadata) |
| `transformer/` | `DiffusionTransformer`: 30 layers, model width 3,840, 30 attention heads, 2 refiner layers | About 6.15 billion |
| `vae/` | `AutoencoderKLQwenImage` with 4 input channels (RGBA) and a 16-channel latent | Not stated |

Source: configuration and safetensors index files in the Hugging Face repositories.[11]

Hugging Face's safetensors metadata reports 6,154,901,056 parameters for the Design repository (BF16) and 6,154,908,736 for the Layer repository (FP32). These numbers equal the byte sizes of the two diffusion-transformer weight files divided by 2 and 4 bytes per parameter respectively, so the "6B" label refers to the diffusion transformer. The complete checkpoint also includes the roughly 17-billion-parameter multimodal language model and the 1.5-billion-parameter connector.[11] The two repositories share the same language-model and VAE configurations. Their diffusion-transformer configurations differ in two fields: `alignment_padding_mode` is `zero_masked` for Design and `learned` for Layer, and `multi_frame_output` is off for Design and on for Layer, which matches the Layer model returning several images.[5][11]

The [vLLM](https://aiwiki.ai/wiki/vllm)-Omni serving recipe for the family calls it a two-stage system. In the first stage, a [Qwen2.5-VL](https://aiwiki.ai/wiki/qwen2_5_vl) vision tower and the 20-layer BailingMoeV2 language model add 256 learned image-query tokens and export their final states, together with hidden states from layers 5, 12 and 20. In the second stage, those conditions go through the connector into "a 30-layer Z-Image DiT with a Qwen-Image RGBA VAE".[13] (See [diffusion transformer](https://aiwiki.ai/wiki/diffusion_transformer).) [ComfyUI](https://aiwiki.ai/wiki/comfyui)'s repackaged files for the models name the text encoder `ming_image_0.1_ling_mini_2.0`.[14]

For the Layer model, the launch post says each layer gets a "Type Token" for its design role. It lists two training measures. "Alpha-Aware Layer Optimization" concentrates on transparent regions and edges. "Composite-Layer Stack Consistency" checks that the layers recombine into the original image.[2]

## Capabilities

### Ming-Image-0.1-Design

Ant Ling says the Design model creates complete UIs, dashboards, infographics and posters "from structured prompts of up to 8K tokens, with consistent layouts, typography, colors, and imagery", and can natively generate transparent-background RGBA assets.[1] The Chinese post names four improvements: expanding long structured prompts of up to 8K; more stable text rendering and layout for titles, buttons, cards and multi-region information; generating a whole design end to end with consistent colour, composition and asset style; and a native RGBA VAE for transparent-background people, products, icons and decorations.[2] The post argues that generating a design in one pass avoids the colour clashes and style mismatches that come from calling a general image model separately for background, illustration and decoration.[2]

The inference code in the companion GitHub repository supports two square output resolutions, 1024 and 2048 pixels, with 2048 x 2048 recommended. The recommended settings are 12 sampling steps, a CFG scale of 1.0 and BF16 precision. The validated hardware is one CUDA GPU with at least 80 GiB of memory.[3][5] The repository recommends running prompts through a separate rewriting step first. An instruction-following vision-language model, either Ling-3.0-flash-VL or qwen3.8-27B, turns a short request into a JSON layout description made of "Figma-style layers" with coordinates, hierarchy and colour values.[5] See [Ling-3.0-flash](https://aiwiki.ai/wiki/ling_3_0_flash). Transparent output is triggered by putting one of ten fixed phrases, five in Chinese and five in English, at the start of the prompt, for example "RGBA, 4-channel, transparent background".[5]

### Ming-Image-0.1-Design-Layer

Ant Ling says the Layer model "converts flattened graphics into 2-9 independently editable RGBA layers". Text, cards, subjects and backgrounds become separate layers that can be moved, replaced, recoloured or reused.[1] The Chinese post says the model targets advertising posters, PPT slides, UI/UX, infographics, e-commerce key visuals and template remixing. It groups layers by design role, keeps transparent edges, fills in occluded regions and tries to make the recombined layers match the original.[2]

The model takes an input image plus a layer specification in the form "Decompose this image into N layers" or "Number of layers: N", or a default request built from a `--num-layers` value. It returns the requested layers plus one leading composite image, which the command-line tool skips when saving the layers as PNG files.[5] The recommended working resolution is 1024 (512 for faster runs), with 12 sampling steps, a CFG scale of 2.0 and BF16 precision, and the output keeps the input image's aspect ratio.[4][5] As with the Design model, the repository describes an optional rewriting step in which a VLM turns a rough layer plan into a detailed one. Text goes in the front layers, a card or banner behind text gets its own layer, the main subject gets its own layer and the background comes last.[5]

## Benchmarks and evaluation

### Artificial Analysis UI/UX Design leaderboard

Artificial Analysis ranks text-to-image models by Elo scores from blind preference votes in its Image Arena, and filters the leaderboard by use case, including a "UIUX Design" category.[7][15] The launch materials used a leaderboard image showing Ming-Image-0.1-Design first in the open-weights view of that category with an Elo score of 1,082. The Chinese post says this was the leaderboard as updated by Artificial Analysis on 18 September 2026.[2][3]

| Snapshot | View | Rank | Elo | Source |
| --- | --- | --- | --- | --- |
| 18 September 2026 (per Ant) | UI/UX Design, open weights | 1 | 1,082 | Launch post and model card image[2][3] |
| 25 September 2026 | UI/UX Design, open weights | 1 | 1,084 (95% CI plus or minus 22; 2,112 samples) | Artificial Analysis[8] |
| 25 September 2026 | UI/UX Design, all models | 16 (range 13-19) | 1,084 | Artificial Analysis[7] |
| 25 September 2026 | All categories, all models | 45 (range 38-50) | 997 | Artificial Analysis[15] |

On 25 September 2026, second place in the open-weights UI/UX view was Ideogram 4.0 (Quality) at 1,047, followed by Ideogram 4.0 at 1,019. The top of the all-model UI/UX view was held by proprietary models led by OpenAI's GPT Image 2.5 Flare (max) at 1,225.[7][8] The Chinese post also reports category win rates of 67.4% for layout, 67.0% for complex composition and 66.7% for text rendering in the UI/UX setting. It also describes vendor-selected head-to-head cases: 12 of 12 wins against Nano Banana 2 Lite, FLUX.2 [max] and Krea 2 Medium on an orchard dashboard prompt, and 10 of 10 wins against [Nano Banana 2](https://aiwiki.ai/wiki/nano_banana_2), Nano Banana Pro and MAI-Image-2.5-Flash on a three-screen recipe app.[2] These figures come from Ant, not from an Artificial Analysis publication.

### Layer decomposition on Crello

The Layer model card reports results on the Crello test set with two metrics, RGB L1 (lower is better) and Alpha soft IoU (higher is better). Each metric is measured at six "Max-allowed Layer Merge" (MLM) settings from 0 to 5, for 12 settings in total.[4] In the table the model appears as "CLEAR-1024 (Ours)" and has the best value in every column. On this basis Ant Ling says the model achieved "the best results across all 12 evaluated settings on the Crello test set".[1][4]

| Method | RGB L1, MLM 0 | RGB L1, MLM 5 | Alpha soft IoU, MLM 0 | Alpha soft IoU, MLM 5 |
| --- | --- | --- | --- | --- |
| VLM Base + Hi-SAM | 0.1197 | 0.0726 | 0.5596 | 0.7589 |
| Yolo Base + Hi-SAM | 0.0962 | 0.0579 | 0.5697 | 0.7897 |
| LayerD | 0.0709 | 0.0396 | 0.7520 | 0.8650 |
| Qwen-Image-Layered-I2L (not open-sourced, fine-tuned on Crello training set) | 0.0594 | 0.0363 | 0.8705 | 0.9160 |
| Qwen-Image-Layered-I2L-640 (open-source release) | 0.1481 | 0.0880 | 0.7131 | 0.8672 |
| Qwen-Image-Layered-I2L-1024 (open-source release) | 0.1409 | 0.0736 | 0.7177 | 0.8749 |
| Ming-Image-0.1-Design-Layer ("CLEAR-1024") | 0.0574 | 0.0314 | 0.8923 | 0.9424 |

Source: vendor-reported table on the Hugging Face model card (selected columns).[4]

The Qwen baseline is [Qwen-Image-Layered](https://aiwiki.ai/wiki/qwen_image), described in arXiv 2512.15603 as an end-to-end diffusion model that decomposes an RGB image into semantically disentangled RGBA layers.[16] The Chinese post says that on the same hardware and software and under a unified setup, the 6B Layer model outperformed the official 20B open-source Qwen version in quality and cut the time for one inference from 795 seconds to 183 seconds, which it rounds to 4.3 times faster.[2] It adds that, after engineering optimization, its end-to-end service at 1024 resolution usually returns a full result in about 20 seconds.[2]

### Reported internal use

The Chinese post gives internal figures from Ant Group products, based on what it calls the current internal testing standard. In Huamei, the time to get editable layers for the first time fell from over 5 minutes to about 40 seconds, which it describes as about 7.5 times faster. The post also says that doing text-replacement edits with GPT-image2 had cost 7 times as much as the Ming-Image model.[2] Lingguang's flash apps that use the models were in internal testing at the time of the announcement.[2]

## Agent Skills

The two skills are published in inclusionAI's `ling-cookbook` repository on GitHub, under `resources/recommended-skills/`, and each carries an MIT License with the notice "Copyright (c) 2026 Ant Group".[17][18] (See [Agent Skills](https://aiwiki.ai/wiki/agent_skills).)

| Skill | What it does | Model service |
| --- | --- | --- |
| Ling UI Design (`ling-ui-design`) | A "visual-first" workflow for text-to-code and screenshot-to-code. The agent chooses a visual direction, generates and checks a reference image, decomposes it into layers, extracts and reviews assets, implements the page, then renders it at target viewports and compares it with the reference | Defaults to Novita's OpenAI-compatible API with model IDs `ming-image-0.1-design` and `ming-image-0.1-design-layer`; other compatible services can be configured |
| Image to Editable PPT (`image-to-editable-ppt`) | Recreates one slide image, screenshot or AI-rendered page as one editable PowerPoint slide: ordinary text becomes text boxes, simple containers become native shapes and complex artwork becomes tightly cropped images | Uses Novita's Ming Image Layer Decoupling API with `ming-image-0.1-design-layer` |

Sources: skill READMEs.[17][18]

The Ling UI Design installer follows the `.agents/skills` convention by default and has an option for [Claude Code](https://aiwiki.ai/wiki/claude_code)'s `.claude/skills` folder.[17] The PPT skill README gives install paths for Codex, Claude Code and [Cursor](https://aiwiki.ai/wiki/cursor), and warns that "The workflow needs a strong model", because the agent has to inspect the image, choose the layer split, accept or reject each layer and write the scene description.[18] Ant Ling summarizes the UI skill's pipeline as "Prompt or screenshot → design → assets → code → browser validation".[1] The Chinese post says the Text-to-Page route shows a complete page proposal in about 15 seconds, compared with about 5 minutes when an agent writes code first and renders afterwards.[2]

## Availability

- **Weights:** Hugging Face (`inclusionAI/Ming-Image-0.1-Design`, `inclusionAI/Ming-Image-0.1-Design-Layer`) and ModelScope, under the MIT License. The model cards list `library_name: custom` and point to the companion `inclusionAI/Ming-Image` GitHub repository, also MIT-licensed, for installation and inference.[3][4][5][6]
- **Serving:** the model cards recommend vLLM-Omni. Its recipe serves both checkpoints through an OpenAI-compatible chat-completions endpoint and lists a test environment of two H100 80GB GPUs, with a single H100 "to be validated". The model cards themselves call one 80 GiB GPU the validated configuration.[3][13]
- **Hosted APIs:** the Chinese post announced two weeks of free API calls for both models on [OpenRouter](https://aiwiki.ai/wiki/openrouter).[2] On 25 September 2026 OpenRouter listed Ming Image 0.1 Design as free, served by a single provider, NovitaAI.[12] Novita documents the Layer model under its OpenAI Image edits protocol, returning one image per layer and a JSON description of each layer's category and bounding box.[19] Artificial Analysis listed an API price of $30.0 per 1,000 images for the Design model.[7]
- **Demos:** Hugging Face Spaces for the Design model (`inclusionAI/ming-image-0-1-design-demo`) and the Layer model (`Xiaolong-Wang/Ming-Image-0.1-Design-Layer`) were running as of 25 September 2026.[3][4]

## Reception

IT Home (IT之家) reported the release on 23 September 2026 under a headline saying the series ranked first among open-source models in the UI design evaluation. The article repeated the official description, the Artificial Analysis score of 1,082 and the limitations Ant itself had listed.[20]

Community ports appeared within days on Hugging Face. They included GGUF quantizations of the Design model (`realrebelai/Ming-Image_GGUFs`, created 23 September) with ComfyUI workflows, and a Comfy-Org repository (created 24 September) that repackages both models for ComfyUI in BF16 and INT8 versions.[14][21] See [GGUF](https://aiwiki.ai/wiki/gguf). As of 25 September, the Design repository had 237 likes on Hugging Face and the Layer repository had 69.[11]

### Limitations stated by the developer

The Chinese post says that the Design model is most reliable at typesetting and text, and less stable on complex hand poses, sequences of actions and fine shadows or reflections. It also says layer decomposition has no single correct answer, since users want different granularities. Heavy stacked light effects, special transparent materials and complex occlusion can leave edge residue or missing layers.[2] Ant said its next steps were better text and complex-layout stability, more controllable layer semantics, better transparency and occlusion completion, and faster inference.[2]

## References

1. Ant Ling (@AntLingAGI), "We're open-sourcing the Ming-Image-0.1-Design family" (thread), X, 22 September 2026. https://x.com/AntLingAGI/status/2102452045374804304
2. 百灵大模型 (Bailing large model team), "Ming-Image-0.1-Design 系列开源，走进真实设计工作流", WeChat, 23 September 2026. https://mp.weixin.qq.com/s/VGdtxfM8kbHIQJw50VD_Sw
3. inclusionAI, "Ming-Image-0.1-Design" model card, Hugging Face. https://huggingface.co/inclusionAI/Ming-Image-0.1-Design
4. inclusionAI, "Ming-Image-0.1-Design-Layer" model card, Hugging Face. https://huggingface.co/inclusionAI/Ming-Image-0.1-Design-Layer
5. inclusionAI, "Ming Image 0.1 Design" (Ming-Image repository README), GitHub. https://github.com/inclusionAI/Ming-Image
6. inclusionAI, "Ming-Image-0.1-Design", ModelScope. https://www.modelscope.cn/models/inclusionAI/Ming-Image-0.1-Design
7. Artificial Analysis, "Text to Image Leaderboard" (category: UIUX Design), accessed 25 September 2026. https://artificialanalysis.ai/image/leaderboard/text-to-image?category=uiux_design
8. Artificial Analysis, "Text to Image Leaderboard" (category: UIUX Design, open weights), accessed 25 September 2026. https://artificialanalysis.ai/image/leaderboard/text-to-image?category=uiux_design&open-weights=true
9. inclusionAI, "INCLUSION AI" (homepage). https://www.inclusion-ai.org/
10. "Ming-Omni: A Unified Multimodal Model for Perception and Generation", arXiv 2506.09344. https://arxiv.org/abs/2506.09344
11. inclusionAI, Ming-Image-0.1-Design and Ming-Image-0.1-Design-Layer repository files (LICENSE, component config.json and safetensors index files) and Hugging Face model API records, accessed 25 September 2026. https://huggingface.co/inclusionAI/Ming-Image-0.1-Design/tree/main ; https://huggingface.co/inclusionAI/Ming-Image-0.1-Design-Layer/tree/main ; https://huggingface.co/api/models/inclusionAI/Ming-Image-0.1-Design
12. OpenRouter, "Ming Image 0.1 Design - API Pricing & Providers", accessed 25 September 2026. https://openrouter.ai/inclusionai/ming-image-0.1-design
13. vLLM-Omni, "Ming-Image 0.1 Design" recipe, GitHub. https://github.com/vllm-project/vllm-omni/blob/main/recipes/inclusionAI/Ming-Image.md
14. Comfy-Org, "Ming-Image" (repackaged model files for ComfyUI), Hugging Face. https://huggingface.co/Comfy-Org/Ming-Image
15. Artificial Analysis, "Text to Image Leaderboard - Top AI Image Models", accessed 25 September 2026. https://artificialanalysis.ai/image/leaderboard/text-to-image
16. Yin et al., "Qwen-Image-Layered: Towards Inherent Editability via Layer Decomposition", arXiv 2512.15603. https://arxiv.org/abs/2512.15603
17. inclusionAI, "Ling UI Design" skill README, ling-cookbook, GitHub. https://github.com/inclusionAI/ling-cookbook/tree/main/resources/recommended-skills/ling-ui-design
18. inclusionAI, "Image to Editable PPT" skill README, ling-cookbook, GitHub. https://github.com/inclusionAI/ling-cookbook/tree/main/resources/recommended-skills/image-to-editable-ppt
19. Novita AI, "Novita AI Ming Image Layer Decoupling API", documentation. https://docs.novita.ai/api-reference/model-apis-ming-image-layer
20. IT之家 (IT Home), "蚂蚁百灵开源 Ming-Image-0.1-Design 系列模型，UI 设计专项评测位列开源第一", 23 September 2026. https://www.ithome.com/1/006/390.htm
21. realrebelai, "Ming-Image_GGUFs", Hugging Face. https://huggingface.co/realrebelai/Ming-Image_GGUFs

