AI Image Generation
AI image generation is the use of artificial intelligence systems to create visual content, including photographs, illustrations, paintings, concept art, and graphic designs, from text descriptions, reference…
Explore Computer Vision through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of Computer Vision.
Showing 1-37 of 37 articles
AI image generation is the use of artificial intelligence systems to create visual content, including photographs, illustrations, paintings, concept art, and graphic designs, from text descriptions, reference…
AI video generation is the use of artificial intelligence systems, predominantly diffusion transformers, to create video clips from text descriptions, still images, or other video, producing sequences of…
Atlas is an early-access world model announced by World Labs on September 1, 2026.
BigGAN is a class-conditional generative adversarial network that, when introduced by DeepMind researchers Andrew Brock, Jeff Donahue, and Karen Simonyan in 2018, set a new state of the art for AI image…
Cartwheel is an American generative AI company that builds 3D character animation tools for games, film and television, advertising, and robotics.
ControlNet is a neural network architecture that adds spatial and structural control to large pretrained text-to-image diffusion models.
CycleGAN (Cycle-Consistent Generative Adversarial Network) is a deep learning architecture for unpaired image-to-image translation.
DCGAN (Deep Convolutional Generative Adversarial Network) is a family of generative adversarial network architectures, introduced in 2015 by Alec Radford, Luke Metz, and Soumith Chintala
A diffusion model is a generative model that learns to transform samples from a simple reference distribution into samples resembling a data distribution by reversing a gradual corruption process.
The Frechet Inception Distance (FID) is the standard metric for measuring the quality of images produced by generative models: it computes the Frechet distance between two multivariate Gaussian distributions…
GAIA-2 (Generative AI for Autonomy 2) is a controllable, multi-camera generative world model for autonomous driving, announced by the British self-driving company Wayve on 26 March 2025.
GAIA-3 is a 15-billion-parameter generative world model for autonomous driving released by Wayve on 2 December 2025, the third generation in the company's GAIA family.
GAIA-4 is a multimodal generative world model for closed-loop autonomous driving simulation, announced by the British self-driving company Wayve on 3 August 2026 as the latest generation of its GAIA family.
Grok Imagine is a generative media product from xAI, the company founded by Elon Musk.
HunyuanWorld 1.0 is an open model from Tencent that generates explorable 3D worlds from a text prompt or a single image.
Ideogram 3.0 is a text-to-image generation model released by Ideogram on March 26, 2025.
A latent diffusion model (LDM) is a type of diffusion model that runs the denoising diffusion process in a compressed latent space learned by a pretrained autoencoder, rather than directly in pixel space…
Luma Dream Machine is a generative AI video and image platform from Luma AI (Luma Labs, Inc.), a San Francisco company, that turns text prompts and still images into short, realistic video clips.
Lumera, short for Light-aware Unified Engine-native Reconstruction and Assembly, is an experimental computer vision benchmark and reference pipeline for turning one RGB image into an editable 3D scene.
Marble is a multimodal generative world model developed by World Labs, the spatial intelligence startup co-founded by Stanford computer scientist Fei-Fei Li.
Nano Banana is the codename, later turned official brand, for Google's native image generation and editing models built into the Gemini ecosystem and developed by Google DeepMind.
Artificial intelligence in photography covers a wide span of techniques, from the computational pipelines baked into modern smartphones to the generative editing tools now built into Photoshop and Lightroom…
Pika 2.5 is a generative video model developed by Pika Labs, the San Francisco based AI video startup co-founded in April 2023 by Stanford AI Lab dropouts Demi Guo and Chenlin Meng.
Runway Act-Two is a generative motion capture and character animation model developed by Runway, publicly introduced on July 15, 2025.
Runway Aleph is an in-context AI video editing model from Runway that transforms and edits an existing video clip from a plain-text instruction, performing a wide range of tasks in a single model: adding…
Sand.ai is an artificial-intelligence company that develops video-generation models, research software, and commercial creation tools.
Seedance is the family of foundation video generation models built by the Seed team at ByteDance, the Chinese internet company that owns TikTok and Douyin.
Seedance 2.5 is a proprietary AI video generation model released by ByteDance on July 31, 2026. It jointly generates audio and video from combinations of text, image, video, and audio inputs.
Seedream is a series of text-to-image and image-editing foundation models built by the Seed research team at ByteDance, the company behind TikTok and Douyin.
Spatial intelligence is the ability of an AI system to perceive, understand, reason about, generate, and interact with three-dimensional space rather than just text or two-dimensional pixels.
StyleGAN is a family of style-based generative adversarial network (GAN) architectures developed by NVIDIA Research for high-quality unconditional image synthesis
Synthesia 3.0 is a major release of the AI video generation platform from Synthesia, the London-based company co-founded in 2017 by Victor Riparbelli, Steffen Tjerrild, Lourdes Agapito and Matthias Niessner.
Wan 2.1-VACE (also written Wan2.1-VACE) is an open-weights video creation and editing model released by Alibaba's Tongyi Lab on May 14, 2025 .
Wan 2.5 is a natively multimodal AI video generation model developed by Alibaba Cloud's Tongyi Lab and previewed at the company's Apsara 2025 conference in Hangzhou on September 24, 2025.
Wayve is a British artificial intelligence company, founded in Cambridge in 2017 by Alex Kendall and Amar Shah, that develops end to end "embodied AI" software for autonomous driving and robotics, and in May…
World Labs is an American spatial-intelligence company, headquartered in San Francisco, that builds "Large World Models" (LWMs), generative AI systems that perceive, generate, reason about and interact with…
pix2pix is a supervised image-to-image translation method that learns a mapping between two visual domains from aligned input-output image pairs.