Generative AI

Explore Generative AI through related topics and the articles other pages reference most.

Explore articles

Reset filters
Browse subtopics: AI Models

Articles that also belong to these categories. Counts cover all of Generative AI.

Showing 1-60 of 64 articles

Amazon Nova Sonic

Amazon Nova Sonic is a real-time speech-to-speech foundation model developed by Amazon and offered through Amazon Bedrock. Announced on April 8, 2025, it is part of the Amazon Nova family of foundation models.

AI Models

Chai-2

Chai-2 is a generative artificial intelligence model from Chai Discovery for designing antibodies and small protein binders from scratch.

AI ModelsAI for Science

ElevenLabs v3

Eleven v3, marketed by ElevenLabs as Eleven v3 (alpha), is a third-generation text-to-speech model that ElevenLabs released in public alpha on June 5, 2025 and described as "the most expressive Text to Speech…

AI ModelsSpeech & Audio AI

FLUX.2

FLUX.2 is the second-generation image generation and editing model family developed by Black Forest Labs, released on November 25, 2025.

AI ModelsDiffusion Models

GAIA-2 (Wayve)

GAIA-2 (Generative AI for Autonomy 2) is a controllable, multi-camera generative world model for autonomous driving, announced by the British self-driving company Wayve on 26 March 2025.

AI ModelsAutonomous Vehicles

GAIA-4 (Wayve)

GAIA-4 is a multimodal generative world model for closed-loop autonomous driving simulation, announced by the British self-driving company Wayve on 3 August 2026 as the latest generation of its GAIA family.

AI ModelsAutonomous Vehicles

GPT Image 1

GPT Image 1 (API identifier gpt-image-1) is a natively multimodal image generation model developed by OpenAI, integrated into ChatGPT on March 25, 2025, and released as a standalone API on April 23, 2025.

AI ModelsImage Generation

Gemini 3.6 Flash

Gemini 3.6 Flash is a proprietary, multimodal large language model released by Google on July 21, 2026. It belongs to the Gemini 3 series and uses the stable API identifier gemini-3.6-flash.

AI ModelsLarge Language Models

Gemini 3.8 Flash

Gemini 3.8 Flash is a multimodal model in Google's Gemini family. Google DeepMind released it on September 2, 2026 as a generally available model for software engineering, tool-using agents, and knowledge work.

AI ModelsGoogle DeepMind

Genie 3

Genie 3 is a general-purpose foundation world model developed by Google DeepMind and announced on August 5, 2025, that generates interactive, navigable 3D environments from a single text prompt and runs in…

AI ModelsGoogle DeepMind

Grok 4.5

Grok 4.5 is a proprietary multimodal large language model and reasoning model in the Grok family. It was developed by SpaceXAI in collaboration with Cursor and released through the xAI API on July 8, 2026.

AI ModelsLarge Language Models

Grok 4.6

Grok 4.6 is a proprietary large language model and reasoning model in the Grok family, developed by SpaceXAI and released jointly with Cursor through the xAI API on August 12, 2026.

AI ModelsLarge Language Models

Hedra Character

Hedra Character is a family of generative video foundation models that turn a single image plus an audio clip (and, in later versions, a text prompt) into a video in which the pictured person or character…

AI ModelsVideo Generation

Hunyuan 3D

Hunyuan 3D is a family of open weight generative artificial intelligence models from Tencent that turn text prompts, single images, sketches, and other inputs into ready to use three dimensional assets…

AI ModelsChinese AI

Imagen 3

Imagen 3 is a text-to-image generation model developed by Google DeepMind, announced at Google I/O on May 14, 2024 and progressively rolled out to users through mid-2024 and into 2025.

AI ModelsGoogle

Imagen 4

Imagen 4 is the fourth-generation text-to-image model developed by Google DeepMind, announced on May 20, 2025, at Google I/O 2025.

AI ModelsGoogle

Luma Dream Machine

Luma Dream Machine is a generative AI video and image platform from Luma AI (Luma Labs, Inc.), a San Francisco company, that turns text prompts and still images into short, realistic video clips.

AI ModelsComputer Vision

Lyria

Lyria is a family of AI music generation models developed by Google DeepMind, spanning text-to-music synthesis, real-time interactive music performance, and full-length song composition.

AI ModelsGoogle DeepMind

MAI-Voice-1

MAI-Voice-1 is a text-to-speech (speech generation) model developed by Microsoft AI, the consumer artificial-intelligence division of Microsoft led by Mustafa Suleyman.

AI Models

Meshy 6

Meshy 6 is the sixth major release of Meshy AI's generative 3D generation platform, a hosted generative AI service that turns text prompts or images into textured 3D models (meshes) for games, animation, and…

AI Models

Midjourney V7

Midjourney V7 is the seventh major text-to-image generation model developed by Midjourney Inc. It launched in alpha on April 3, 2025, and became the platform's default model on June 17, 2025.

AI ModelsImage Generation

NVIDIA Cosmos

NVIDIA Cosmos is a world foundation model platform developed by NVIDIA for physical AI applications, including autonomous vehicles and robotics.

AI ModelsEmbodied AI

NVIDIA Picasso

NVIDIA Picasso is a cloud-based generative AI foundry from NVIDIA for building, training, and deploying visual generative models that produce images, video, and 3D content from text prompts.

AI HardwareAI Inference

Nano Banana

Nano Banana is the codename, later turned official brand, for Google's native image generation and editing models built into the Gemini ecosystem and developed by Google DeepMind.

AI ModelsComputer Vision

Pika (video generation)

Pika is an artificial intelligence video generation platform developed by Pika Labs, Inc. that lets users create and edit short videos from text prompts, images, and existing clips, and it is best known for…

AI CompaniesAI Models

Pika 2.5

Pika 2.5 is a generative video model developed by Pika Labs, the San Francisco based AI video startup co-founded in April 2023 by Stanford AI Lab dropouts Demi Guo and Chenlin Meng.

AI ModelsComputer Vision

Qwen-Image

Qwen-Image is an open-weight image-generation foundation model released by Alibaba's Qwen team in August 2025.

AI Models

Qwen-Image-3.0

Qwen-Image-3.0 is a text-to-image foundation model announced by Alibaba's Qwen team on July 21, 2026, as the third generation of the Qwen-Image series .

AI ModelsChinese AI

Recraft V3

Recraft V3 is a text-to-image generation model developed by Recraft AI and released on October 30, 2024, that became the first model to reach the number-one position on the Artificial Analysis Text-to-Image…

AI ModelsImage Generation

Rodin Gen-2

Rodin Gen-2 is a generative artificial intelligence model for producing three dimensional assets from text prompts or reference images, developed by the Shanghai based company Deemos and delivered through the…

AI ModelsChinese AI

Runway Aleph

Runway Aleph is an in-context AI video editing model from Runway that transforms and edits an existing video clip from a plain-text instruction, performing a wide range of tasks in a single model: adding…

AI ModelsComputer Vision

Seedance

Seedance is the family of foundation video generation models built by the Seed team at ByteDance, the Chinese internet company that owns TikTok and Douyin.

AI ModelsChinese AI

Seedance 2.5

Seedance 2.5 is a proprietary AI video generation model released by ByteDance on July 31, 2026. It jointly generates audio and video from combinations of text, image, video, and audio inputs.

AI ModelsChinese AI

Seedream

Seedream is a series of text-to-image and image-editing foundation models built by the Seed research team at ByteDance, the company behind TikTok and Douyin.

AI ModelsChinese AI

Sesame CSM

Sesame CSM (Conversational Speech Model) is an open weights speech generation model from Sesame AI, a San Francisco startup co-founded by former Oculus chief executive Brendan Iribe.

AI ModelsOpen Source AI

Shap-E

Shap-E is a conditional generative model for 3D assets developed by OpenAI, introduced in the paper "Shap-E: Generating Conditional 3D Implicit Functions" by Heewoo Jun and Alex Nichol, submitted to arXiv on…

AI ModelsOpenAI

Sora 2

Sora 2 is a text-to-video and audio generation model developed by OpenAI, released on September 30, 2025, that OpenAI called "the GPT-3.5 moment for video." It succeeded the original Sora research preview from…

AI ModelsMultimodal AI

Suno v5

Suno v5 is the fifth-generation AI music generation model from Suno Inc., the Cambridge, Massachusetts startup, released on September 23, 2025 to Pro and Premier subscribers as what Suno called "the world's…

AI ModelsMusic & Audio Generation

Synthesia 3.0

Synthesia 3.0 is a major release of the AI video generation platform from Synthesia, the London-based company co-founded in 2017 by Victor Riparbelli, Steffen Tjerrild, Lourdes Agapito and Matthias Niessner.

AI ModelsComputer Vision