GPT Image 1
GPT Image 1 (API identifier gpt-image-1) is a natively multimodal image generation model developed by OpenAI, integrated into ChatGPT on March 25, 2025, and released as a standalone API on April 23, 2025.
Explore Multimodal AI through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of Multimodal AI.
Showing 1-12 of 12 articles
GPT Image 1 (API identifier gpt-image-1) is a natively multimodal image generation model developed by OpenAI, integrated into ChatGPT on March 25, 2025, and released as a standalone API on April 23, 2025.
Gemini 3.6 Flash is a proprietary, multimodal large language model released by Google on July 21, 2026. It belongs to the Gemini 3 series and uses the stable API identifier gemini-3.6-flash.
Gemini 3.7 Flash is a proprietary, multimodal large language model released by Google on August 13, 2026.
Gemini 3.8 Flash is a multimodal model in Google's Gemini family. Google DeepMind released it on September 2, 2026 as a generally available model for software engineering, tool-using agents, and knowledge work.
Gemini Omni is a family of proprietary multimodal AI models from Google DeepMind for generating and editing media.
MAGI-2 Preview is a public research release of a unified audio-video generation model developed by Sand.ai.
Muse Image is a proprietary image-generation and editing system developed by Meta Superintelligence Labs, a division of Meta.
Nano Banana Pro is a professional grade image generation and editing model from Google DeepMind, released on November 20, 2025.
Pika is an artificial intelligence video generation platform developed by Pika Labs, Inc. that lets users create and edit short videos from text prompts, images, and existing clips, and it is best known for…
Seedream 4.0 is a unified image generation and editing model built by the Seed team at ByteDance.
Sora 2 is a text-to-video and audio generation model developed by OpenAI, released on September 30, 2025, that OpenAI called "the GPT-3.5 moment for video." It succeeded the original Sora research preview from…
Text-to-image models are generative artificial intelligence systems that synthesize a new image from a natural-language description, called a prompt.