AI Model Release Timeline (2022-2026)
The pace of frontier AI model releases went from a single landmark launch in late 2022 to a new flagship roughly every one to two weeks by 2026.
Explore AI Models through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of AI Models.
Showing 1-60 of 63 articles
The pace of frontier AI model releases went from a single landmark launch in late 2022 to a new flagship roughly every one to two weeks by 2026.
Amazon Nova Sonic is a real-time speech-to-speech foundation model developed by Amazon and offered through Amazon Bedrock. Announced on April 8, 2025, it is part of the Amazon Nova family of foundation models.
Seed3D 2.0 is a 3D-asset generation model released by ByteDance's Seed research team on April 23, 2026.
Chai-2 is a generative artificial intelligence model from Chai Discovery for designing antibodies and small protein binders from scratch.
ElevenLabs Music (also marketed as Eleven Music) is an AI music generation product developed by ElevenLabs, the voice AI company founded in 2022 by Piotr Dabkowski and Mati Staniszewski.
Eleven v3, marketed by ElevenLabs as Eleven v3 (alpha), is a third-generation text-to-speech model that ElevenLabs released in public alpha on June 5, 2025 and described as "the most expressive Text to Speech…
FLUX 3 Video is a video generation model by Black Forest Labs (BFL), released into general availability on August 4, 2026.
FLUX.1 is a family of text-to-image generation models developed by Black Forest Labs, released on August 1, 2024.
FLUX.2 is the second-generation image generation and editing model family developed by Black Forest Labs, released on November 25, 2025.
GAIA-2 (Generative AI for Autonomy 2) is a controllable, multi-camera generative world model for autonomous driving, announced by the British self-driving company Wayve on 26 March 2025.
GAIA-3 is a 15-billion-parameter generative world model for autonomous driving released by Wayve on 2 December 2025, the third generation in the company's GAIA family.
GAIA-4 is a multimodal generative world model for closed-loop autonomous driving simulation, announced by the British self-driving company Wayve on 3 August 2026 as the latest generation of its GAIA family.
GPT Image 1 (API identifier gpt-image-1) is a natively multimodal image generation model developed by OpenAI, integrated into ChatGPT on March 25, 2025, and released as a standalone API on April 23, 2025.
OpenAI's GPT lineage runs from GPT-1 (June 2018, 117 million parameters, a 512-token context window) to the GPT-5.6 "Sol, Terra, Luna" family that entered limited preview on June 26, 2026.
Gemini 3.6 Flash is a proprietary, multimodal large language model released by Google on July 21, 2026. It belongs to the Gemini 3 series and uses the stable API identifier gemini-3.6-flash.
Gemini 3.7 Flash is a proprietary, multimodal large language model released by Google on August 13, 2026.
Gemini 3.8 Flash is a multimodal model in Google's Gemini family. Google DeepMind released it on September 2, 2026 as a generally available model for software engineering, tool-using agents, and knowledge work.
Gemini Omni is a family of proprietary multimodal AI models from Google DeepMind for generating and editing media.
Genie 3 is a general-purpose foundation world model developed by Google DeepMind and announced on August 5, 2025, that generates interactive, navigable 3D environments from a single text prompt and runs in…
Grok 4.5 is a proprietary multimodal large language model and reasoning model in the Grok family. It was developed by SpaceXAI in collaboration with Cursor and released through the xAI API on July 8, 2026.
Grok 4.6 is a proprietary large language model and reasoning model in the Grok family, developed by SpaceXAI and released jointly with Cursor through the xAI API on August 12, 2026.
Grok Imagine is a generative media product from xAI, the company founded by Elon Musk.
Hedra Character is a family of generative video foundation models that turn a single image plus an audio clip (and, in later versions, a text prompt) into a video in which the pictured person or character…
Avatar IV is the fourth generation of the AI avatar engine from HeyGen, the AI video company co-founded in 2020 by Joshua Xu and Wayne Liang.
Hume Octave 2 is a multilingual emotional text-to-speech model released by Hume AI on October 1, 2025.
Hunyuan 3D is a family of open weight generative artificial intelligence models from Tencent that turn text prompts, single images, sketches, and other inputs into ready to use three dimensional assets…
Ideogram 3.0 is a text-to-image generation model released by Ideogram on March 26, 2025.
Imagen 3 is a text-to-image generation model developed by Google DeepMind, announced at Google I/O on May 14, 2024 and progressively rolled out to users through mid-2024 and into 2025.
Imagen 4 is the fourth-generation text-to-image model developed by Google DeepMind, announced on May 20, 2025, at Google I/O 2025.
Luma Dream Machine is a generative AI video and image platform from Luma AI (Luma Labs, Inc.), a San Francisco company, that turns text prompts and still images into short, realistic video clips.
Lyria is a family of AI music generation models developed by Google DeepMind, spanning text-to-music synthesis, real-time interactive music performance, and full-length song composition.
MAGI-2 Preview is a public research release of a unified audio-video generation model developed by Sand.ai.
MAI-Voice-1 is a text-to-speech (speech generation) model developed by Microsoft AI, the consumer artificial-intelligence division of Microsoft led by Mustafa Suleyman.
Meshy 6 is the sixth major release of Meshy AI's generative 3D generation platform, a hosted generative AI service that turns text prompts or images into textured 3D models (meshes) for games, animation, and…
Midjourney V7 is the seventh major text-to-image generation model developed by Midjourney Inc. It launched in alpha on April 3, 2025, and became the platform's default model on June 17, 2025.
Muse Image is a proprietary image-generation and editing system developed by Meta Superintelligence Labs, a division of Meta.
NVIDIA Cosmos is a world foundation model platform developed by NVIDIA for physical AI applications, including autonomous vehicles and robotics.
NVIDIA Picasso is a cloud-based generative AI foundry from NVIDIA for building, training, and deploying visual generative models that produce images, video, and 3D content from text prompts.
Nano Banana is the codename, later turned official brand, for Google's native image generation and editing models built into the Gemini ecosystem and developed by Google DeepMind.
Nano Banana 2 Lite is a text-to-image generation and image editing model released by Google DeepMind on June 30, 2026.
Pika is an artificial intelligence video generation platform developed by Pika Labs, Inc. that lets users create and edit short videos from text prompts, images, and existing clips, and it is best known for…
Pika 2.5 is a generative video model developed by Pika Labs, the San Francisco based AI video startup co-founded in April 2023 by Stanford AI Lab dropouts Demi Guo and Chenlin Meng.
Qwen-Image is an open-weight image-generation foundation model released by Alibaba's Qwen team in August 2025.
Qwen-Image-3.0 is a text-to-image foundation model announced by Alibaba's Qwen team on July 21, 2026, as the third generation of the Qwen-Image series .
Recraft V3 is a text-to-image generation model developed by Recraft AI and released on October 30, 2024, that became the first model to reach the number-one position on the Artificial Analysis Text-to-Image…
Rodin Gen-2 is a generative artificial intelligence model for producing three dimensional assets from text prompts or reference images, developed by the Shanghai based company Deemos and delivered through the…
Runway Act-Two is a generative motion capture and character animation model developed by Runway, publicly introduced on July 15, 2025.
Runway Aleph is an in-context AI video editing model from Runway that transforms and edits an existing video clip from a plain-text instruction, performing a wide range of tasks in a single model: adding…
Runway Gen-4 is the fourth-generation video generation model developed by Runway (company), announced and released on March 31, 2025.
Seedance is the family of foundation video generation models built by the Seed team at ByteDance, the Chinese internet company that owns TikTok and Douyin.
Seedance 2.5 is a proprietary AI video generation model released by ByteDance on July 31, 2026. It jointly generates audio and video from combinations of text, image, video, and audio inputs.
Seedream is a series of text-to-image and image-editing foundation models built by the Seed research team at ByteDance, the company behind TikTok and Douyin.
Sesame CSM (Conversational Speech Model) is an open weights speech generation model from Sesame AI, a San Francisco startup co-founded by former Oculus chief executive Brendan Iribe.
Shap-E is a conditional generative model for 3D assets developed by OpenAI, introduced in the paper "Shap-E: Generating Conditional 3D Implicit Functions" by Heewoo Jun and Alex Nichol, submitted to arXiv on…
Sora 2 is a text-to-video and audio generation model developed by OpenAI, released on September 30, 2025, that OpenAI called "the GPT-3.5 moment for video." It succeeded the original Sora research preview from…
Stable Audio 2.5 is an enterprise focused text-to-audio generation model released by Stability AI on September 10, 2025.
Suno v5 is the fifth-generation AI music generation model from Suno Inc., the Cambridge, Massachusetts startup, released on September 23, 2025 to Pro and Premier subscribers as what Suno called "the world's…
Synthesia 3.0 is a major release of the AI video generation platform from Synthesia, the London-based company co-founded in 2017 by Victor Riparbelli, Steffen Tjerrild, Lourdes Agapito and Matthias Niessner.
Text-to-image models are generative artificial intelligence systems that synthesize a new image from a natural-language description, called a prompt.
Tripo P1, marketed in full as Tripo Smart Mesh P1.0, is a production grade native 3D diffusion model that generates clean, engine ready 3D meshes from text or image prompts in as little as two seconds.