Generative AI

Explore Generative AI through related topics and the articles other pages reference most.

Explore articles

Browse subtopics (61)

Articles that also belong to these categories. Counts cover all of Generative AI.

Showing 121-180 of 264 articles

IP-Adapter

IP-Adapter (short for Image Prompt Adapter) is a lightweight neural network module that adds image-prompt conditioning to a pretrained text-to-image diffusion model, allowing a reference image to guide…

Deep Learning

Ian Goodfellow

Ian Goodfellow is an American computer scientist and machine learning researcher best known for inventing the generative adversarial network (GAN) in 2014 and for being the lead author of the textbook Deep…

Deep LearningPeople

Ideogram

Ideogram is a Toronto-based artificial intelligence company, founded in 2022 by four former Google Brain researchers who built Google's Imagen system

AI CompaniesImage Generation

Imagen 3

Imagen 3 is a text-to-image generation model developed by Google DeepMind, announced at Google I/O on May 14, 2024 and progressively rolled out to users through mid-2024 and into 2025.

AI ModelsGoogle

Imagen 4

Imagen 4 is the fourth-generation text-to-image model developed by Google DeepMind, announced on May 20, 2025, at Google I/O 2025.

AI ModelsGoogle

Jonathan Ho

Jonathan Ho is a machine learning researcher best known as the lead author of "Denoising Diffusion Probabilistic Models" (DDPM), the 2020 paper that made diffusion models practical for high quality image…

Deep LearningPeople

Kling 2.1

Kling 2.1 is a video generation model developed by Kuaishou Technology, the Beijing-based internet and short-video company behind China's second-largest short-video platform.

Chinese AIVideo Generation

Kling 3.0

Kling 3.0 is the third-generation AI video generation model family released by Kuaishou on February 4, 2026, comprising four models (Video 3.0, Video 3.0 Omni, Image 3.0, and Image 3.0 Omni) that generate…

Chinese AIVideo Generation

Krea AI

Krea AI (legally Krea, krea.ai) is a San Francisco generative AI company, founded in 2022, that builds a browser-based creative suite for image generation, video creation, three-dimensional asset production…

AI CompaniesImage Generation

LAION-5B

LAION-5B is an open dataset of approximately 5.85 billion CLIP-filtered image and text pairs scraped from the public internet, released by LAION (Large-scale Artificial Intelligence Open Network) on March 31

Data & Datasets

Latent Space

A latent space is the vector space a machine learning model maps its inputs into, where each input becomes a point (a latent vector or latent code) and the geometry of the space carries information the raw…

Deep LearningInterpretability

Latent diffusion model

A latent diffusion model (LDM) is a type of diffusion model that runs the denoising diffusion process in a compressed latent space learned by a pretrained autoencoder, rather than directly in pixel space…

Computer VisionDeep Learning

Leonardo.AI

Leonardo.AI (stylized as Leonardo.Ai) is a Sydney-based generative AI platform for image generation, video creation, and creative design, founded in December 2022 by CEO JJ Fiasson and five co-founders and…

AI CompaniesImage Generation

Luma AI

Luma AI (operating as Luma Labs, Inc.) is a generative AI company headquartered in Palo Alto, California, that builds text-to-video, 3D, and image models, best known for its Dream Machine platform and its Ray…

AI CompaniesVideo Generation

Luma Dream Machine

Luma Dream Machine is a generative AI video and image platform from Luma AI (Luma Labs, Inc.), a San Francisco company, that turns text prompts and still images into short, realistic video clips.

AI ModelsComputer Vision

Lumera

Lumera, short for Light-aware Unified Engine-native Reconstruction and Assembly, is an experimental computer vision benchmark and reference pipeline for turning one RGB image into an editable 3D scene.

Computer Vision

Lyria

Lyria is a family of AI music generation models developed by Google DeepMind, spanning text-to-music synthesis, real-time interactive music performance, and full-length song composition.

AI ModelsGoogle DeepMind

MAI-Voice-1

MAI-Voice-1 is a text-to-speech (speech generation) model developed by Microsoft AI, the consumer artificial-intelligence division of Microsoft led by Mustafa Suleyman.

AI Models

MaskGIT

MaskGIT, short for Masked Generative Image Transformer, is an image-synthesis method introduced by Google Research in the 2022 paper "MaskGIT: Masked Generative Image Transformer" by Huiwen Chang, Han Zhang…

Deep Learning

Masked Autoregressive (MAR) generation

Masked Autoregressive (MAR) generation is an image-generation method introduced in the 2024 paper "Autoregressive Image Generation without Vector Quantization" by Tianhong Li, Yonglong Tian, He Li, Mingyang…

Deep Learning

Meshy 6

Meshy 6 is the sixth major release of Meshy AI's generative 3D generation platform, a hosted generative AI service that turns text prompts or images into textured 3D models (meshes) for games, animation, and…

AI Models

Midjourney V7

Midjourney V7 is the seventh major text-to-image generation model developed by Midjourney Inc. It launched in alpha on April 3, 2025, and became the platform's default model on June 17, 2025.

AI ModelsImage Generation

Movie Gen

Movie Gen is a suite of media-generation foundation models from Meta AI, announced on October 4, 2024, that generates high-definition video with synchronized audio from text prompts.

Meta AIVideo Generation

Music

AI in music is the use of artificial intelligence, especially machine learning and generative AI, to compose, perform, mix, master, transcribe, voice-clone, and reproduce music.

AI Tools & ProductsSpeech & Audio AI

MusicGen

MusicGen is an open-weights text-to-music generation model from Meta AI's Fundamental AI Research (FAIR) team, released on 8 June 2023, that generates roughly 30-second music clips from a text prompt, a…

Meta AIMusic & Audio Generation

MusicLM

MusicLM is a text-to-music generation model from Google Research that generates high-fidelity music at 24 kHz from natural language descriptions and keeps that audio consistent over several minutes.

GoogleMusic & Audio Generation

NVIDIA Cosmos

NVIDIA Cosmos is a world foundation model platform developed by NVIDIA for physical AI applications, including autonomous vehicles and robotics.

AI ModelsEmbodied AI

NVIDIA Cosmos 3

NVIDIA Cosmos 3 is an open family of "world foundation models" for physical AI that NVIDIA launched on June 1, 2026 at GTC Taipei, held alongside COMPUTEX 2026.

NVIDIARobotics

NVIDIA Picasso

NVIDIA Picasso is a cloud-based generative AI foundry from NVIDIA for building, training, and deploying visual generative models that produce images, video, and 3D content from text prompts.

AI HardwareAI Inference

Nano Banana

Nano Banana is the codename, later turned official brand, for Google's native image generation and editing models built into the Gemini ecosystem and developed by Google DeepMind.

AI ModelsComputer Vision

OpusClip

OpusClip is an AI video-repurposing tool that turns long-form videos into short, social-ready clips, ranking the moments most likely to go viral and adding animated captions, automatic reframing, B-roll, and…

AI CompaniesVideo Generation

Parti (text-to-image model)

Parti (Pathways Autoregressive Text-to-Image) is a text-to-image generation model from Google Research that produces images from natural-language descriptions by treating the task as a sequence-to-sequence…

GoogleImage Generation

Perplexity Comet

Comet is an AI-powered, agentic web browser developed by Perplexity AI. It is built on the Chromium open-source project and ships with an integrated AI assistant that can summarise pages, answer questions…

AI AgentsAI Companies

Perplexity Finance

Perplexity Finance is a dedicated financial research product from Perplexity AI, available at perplexity.ai/finance, that combines real-time market data, company filings, earnings call transcripts, and…

AI CompaniesFinance AI

Phind

Phind was an AI-powered answer engine built for software developers. The product combined a live web index with fine-tuned large language models to return cited, code-aware answers to programming questions…

AI CompaniesDeveloper Tools