Muse Image
Last edited
Fact-checked
In review queue
Sources
6 citations
Revision
v1 · 1,945 words
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Muse Image is a proprietary image-generation and editing system developed by Meta Superintelligence Labs, a division of Meta. Meta launched it on July 7, 2026 as the lab's first released image model and made it the image generator in Meta AI. It accepts text, one or more reference images, and visual edit annotations, then produces or revises still images.[1][2]
Unlike a simple prompt-to-image service, Muse Image can operate through an agentic workflow. It may plan, search the web, write and execute code, inspect a draft, and refine its output. It can also plan jointly and share tools with the separate Muse Spark reasoning model. Meta disclosed these system behaviors and several evaluation results, but not the underlying generator architecture, parameter count, training corpus, public weights, or a Muse Image developer API.[1][3]
Release and availability
Meta released Muse Image in the Meta AI app and on meta.ai. The initial product rollout also covered more than 30 effects for Instagram Stories in the United States and image generation in direct chats with Meta AI on WhatsApp in limited countries. Facebook, Messenger, more Instagram and WhatsApp surfaces, additional countries, and Meta Advantage+ access for advertisers were announced as future expansions rather than launch availability.[1][2]
| Attribute | Publicly disclosed information |
|---|---|
| Developer | Meta Superintelligence Labs, part of Meta |
| Release date | July 7, 2026 |
| Model type | Proprietary image generation and editing system |
| Inputs | Conversational text, a single image, multiple reference images, and visual annotations or sketches |
| Output | Generated or edited still images |
| Launch access | Meta AI app, meta.ai, Instagram Stories in the US, and WhatsApp in limited countries |
| Parameter count | Not disclosed |
| Public weights | None identified |
| Muse Image API | None announced at launch |
| Model license | Not disclosed |
Meta described normal Meta AI use as free for "everyday creation" and said heavier use was available through Meta subscription plans. Its announcements did not define the free allowance, identify a Muse Image-specific subscription tier, publish a numerical rate limit, or give a per-image price.[2] Axios reported on launch day that Meta was still evaluating whether to make Muse Image available to outside developers.[5]
Generation and editing
Muse Image supports both text-to-image generation and instruction-based editing. A user can start with a text description, submit an existing photo, or interleave text with several reference images. Meta says the multi-reference mode can combine people, objects, clothing, styles, and environments in a single composition.[1]
Editing can continue over multiple conversational turns while Meta AI retains the thread context. In the product interface, a user can also circle, sketch, or annotate a region to indicate a requested change. Meta's examples include removing a photobomber, restoring an old photograph, changing a hairstyle, transferring a style, combining a selfie with a vacation image, and redesigning a room with products found on the web or Facebook Marketplace.[2]
The system is also intended to handle image text and structured graphics. Meta demonstrated prompts for infographics, plots, and functional QR codes. During reinforcement learning, Muse Image learned to write and run code that creates a precise plot or QR code, then condition image generation on the rendered result. This is a code-assisted workflow, not evidence that the image model natively computes every graphical element without tools.[1]
Muse Image and Muse Spark can combine code and media generation for larger outputs such as animated GIFs, websites with embedded images, and interactive visual games. Those examples describe a multi-model, tool-using system. They do not mean that Muse Image by itself returns website source code, a game, or video as its native output.[1]
Agentic workflow
Meta presents Muse Image as an agent that can choose among generation, local editing, regeneration, search, and code execution. Web search supplies current information and visual references for prompts involving real-world facts. An internal ablation reported a higher win rate when search was enabled on knowledge-intensive prompts, although Meta did not publish the full prompt set, sample count, raw generations, or a complete scoring protocol.[1]
Self-refinement is another part of the workflow. Meta says the behavior emerged during reinforcement learning because revising a weak image produced higher reward. Depending on the problem, the system can make a small local correction, generate a new image, or try a tool. The release does not identify the reinforcement-learning algorithm, reward model, training annotators, or mixture of generated and human feedback.[1]
Meta also reported that increasing inference-time compute caused more reasoning, tool calls, and refinement steps, with approximately log-linear improvement in an internal human-preference Elo measure. In Meta's comparison, deliberate reasoning continued to improve after best-of-N sampling, which generates several candidates and selects one, began to saturate. The computation combines text tokens used for reasoning and visual tokens used for image generation. These details describe inference behavior, but do not reveal whether the underlying image generator uses diffusion, autoregression, flow matching, or another architecture.[1]
Architecture and training disclosure
The public launch identifies Muse Image as a multimodal AI system and explicitly mentions reinforcement learning, reasoning tokens, visual-generation tokens, search, code execution, and tool sharing with Muse Spark. It also says the previewed Muse Video was built on the same pretraining base. Meta did not provide a layer-level architecture, parameter count, tokenizer specification, image encoder or decoder design, context limit, training compute, inference hardware, or data mixture.[1]
No public source accompanying the release identifies the image, text, web, or preference datasets used to train Muse Image. The launch-day product briefly allowed a user to reference a public Instagram account when generating an image, but that was an inference-time feature and is not evidence that Instagram photos formed the training corpus. Meta removed the feature three days later.[2][6]
As of July 24, 2026, Meta's Muse Image materials linked to no technical paper, model card, safety report, public checkpoint, source repository, dataset, or evaluation package. They also did not specify output resolution, supported aspect ratios, file formats, number of images per request, latency, deterministic controls, or commercial-use license terms for generated files.[1][2]
Evaluation
Meta reported that Muse Image ranked second on Arena's text-to-image, single-image edit, and multi-image edit leaderboards on July 5, 2026. The charts placed GPT Image 2 first in all three categories.[1]
| Arena category | Muse Image score and rank | First-place score |
|---|---|---|
| Text-to-image | 1,280 +/- 7, rank 2 | GPT Image 2, 1,385 +/- 5 |
| Single-image edit | 1,405 +/- 6, rank 2 | GPT Image 2, 1,466 +/- 4 |
| Multi-image edit | 1,399 +/- 6, rank 2 | GPT Image 2, 1,454 +/- 5 |
Arena rankings are based on anonymous head-to-head battles in which users choose a preferred output. The service aggregates those pairwise votes with a Bradley-Terry rating system similar to Elo, and model names are revealed after voting. This provides a large-scale measure of human preference rather than a static test with one ground-truth answer. It is also time-sensitive: ranks and confidence intervals can change as new votes and models arrive.[4]
The Arena positions do not by themselves measure factual accuracy, copyright compliance, safety, latency, or cost. Meta's additional search, self-refinement, and test-time-scaling results were internal ablations. The public post labels them as win rates or Elo comparisons but does not release enough evaluation data or code for an independent replication.[1]
Provenance and the Instagram reference rollback
Muse Image includes Content Seal, an invisible watermarking system. Meta says images created in the Meta AI app and on meta.ai carry a hidden provenance signal that remains detectable after cropping, compression, resizing, or a screenshot. It also released a detector as a preview. The launch did not claim that every Muse Image output on Instagram or WhatsApp had the signal, so Content Seal coverage should not be assumed beyond the surfaces Meta named.[1]
A Content Seal result is also narrower than universal AI-generated media detection. A positive check is intended to show that the Meta signal is present. The release supplies no independent robustness test, false-positive or false-negative rate, adversarial-removal evaluation, or cryptographic specification, and an unmarked image is not thereby proven to be human-made.[1]
Muse Image's launch included a separate feature that let users mention public Instagram accounts as visual references. Public-account images were automatically available for that use unless an account owner changed the applicable setting. The feature drew privacy and likeness criticism, including concern from SAG-AFTRA about nonconsensual digital replicas. Meta removed it on July 10, saying it had "missed the mark." The removal affected that reference mechanism, not the entire Muse Image service.[2][6]
Relationship to Muse Spark and Muse Video
Muse Image and Muse Spark are separate models with complementary roles. Muse Spark is a general multimodal reasoning model, while Muse Image generates and edits images. In a combined workflow they can share tools and plan jointly, with Spark helping to orchestrate code, search, and media-generation steps.[1]
This distinction also matters for developer access. Meta opened its Meta Model API public preview for Muse Spark 1.1 on July 9, two days after the image launch. The API announcement specifically names Muse Spark 1.1 and treats Muse Image as a separate release. Spark's API availability therefore does not establish a Muse Image endpoint or developer license.[3]
Muse Video was not released alongside Muse Image. Meta presented it as an early preview that was "coming soon" to creators and Meta AI. The preview shares a pretraining base with Muse Image and is intended to generate video with native audio, but Meta acknowledged current gaps in audio-video synchronization and physically accurate fast motion. Meta also showed a number 3 Arena text-to-video position dated July 5, but that evaluation does not change the product's unreleased preview status.[1]
Limitations
Muse Image's agentic design makes its public product capabilities broader than its model disclosure. Search can supply timely references, code can construct exact graphical elements, Muse Spark can help plan, and repeated inference can refine a draft. Consequently, a successful demonstration does not show which result came from the image model alone and which depended on retrieval, code, another model, or extra compute.[1]
Reproducibility is limited by the absence of weights, code, training data, an architecture specification, a model card, and detailed internal evaluations. Geographic access was uneven at launch, consumer limits and subscription prices were unspecified, and no public developer API was announced. Content Seal offers a provenance signal on two named surfaces, but Meta did not publish an independent robustness evaluation. The rapid withdrawal of the public Instagram reference feature also exposed a consent and product-governance problem that the technical launch did not anticipate.[1][2][5][6]
The dated Arena results provide useful comparative evidence, but they measure preference within Arena's evolving battle distribution. They should not be read as a permanent ranking or as a certification of factual accuracy, rights compliance, or safe use. The disclosed material supports describing Muse Image as a capable released image-generation system, while leaving its architecture, data, licensing, and several safety properties unresolved.[1][4]
References
- Meta Superintelligence Labs. Introducing Muse Image and Muse Video. Meta AI, July 7, 2026. ↩
- Meta. Introducing Muse Image: Image Generation Built for Your World. Meta Newsroom, July 7, 2026; updated July 10, 2026. ↩
- Meta Superintelligence Labs. Introducing Muse Spark 1.1. Meta AI, July 9, 2026. ↩
- Arena Intelligence. Arena FAQ. Accessed July 24, 2026. ↩
- Fried, Ina. Meta's AI catch-up effort gets a new look. Axios, July 7, 2026. ↩
- Veiga, Alex. Amid criticism, Meta reins in new AI tool that automatically accessed public Instagram images. Associated Press, July 11, 2026. ↩
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.