HunyuanWorld 1.0
HunyuanWorld 1.0 is an open model from Tencent that generates explorable 3D worlds from a text prompt or a single image.
Explore Generative AI through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of Generative AI.
Showing 121-180 of 264 articles
HunyuanWorld 1.0 is an open model from Tencent that generates explorable 3D worlds from a text prompt or a single image.
IP-Adapter (short for Image Prompt Adapter) is a lightweight neural network module that adds image-prompt conditioning to a pretrained text-to-image diffusion model, allowing a reference image to guide…
Ian Goodfellow is an American computer scientist and machine learning researcher best known for inventing the generative adversarial network (GAN) in 2014 and for being the lead author of the textbook Deep…
Ideogram is a Toronto-based artificial intelligence company, founded in 2022 by four former Google Brain researchers who built Google's Imagen system
Ideogram 3.0 is a text-to-image generation model released by Ideogram on March 26, 2025.
Imagen is a family of text-to-image diffusion models developed by Google, first introduced in May 2022 and as of 2026 in its fourth generation (Imagen 4).
Imagen 3 is a text-to-image generation model developed by Google DeepMind, announced at Google I/O on May 14, 2024 and progressively rolled out to users through mid-2024 and into 2025.
Imagen 4 is the fourth-generation text-to-image model developed by Google DeepMind, announced on May 20, 2025, at Google I/O 2025.
Jonathan Ho is a machine learning researcher best known as the lead author of "Denoising Diffusion Probabilistic Models" (DDPM), the 2020 paper that made diffusion models practical for high quality image…
Jukebox is a neural network for music generation developed by OpenAI that produces music, including rudimentary singing, as raw audio across a range of genres and artist styles.
Kling is an AI video generation model developed by Kuaishou Technology, a Chinese short-video platform company publicly traded on the Hong Kong Stock Exchange.
Kling 2.1 is a video generation model developed by Kuaishou Technology, the Beijing-based internet and short-video company behind China's second-largest short-video platform.
Kling 3.0 is the third-generation AI video generation model family released by Kuaishou on February 4, 2026, comprising four models (Video 3.0, Video 3.0 Omni, Image 3.0, and Image 3.0 Omni) that generate…
Krea AI (legally Krea, krea.ai) is a San Francisco generative AI company, founded in 2022, that builds a browser-based creative suite for image generation, video creation, three-dimensional asset production…
LAION-5B is an open dataset of approximately 5.85 billion CLIP-filtered image and text pairs scraped from the public internet, released by LAION (Large-scale Artificial Intelligence Open Network) on March 31
Latent Consistency Models (LCMs) are a family of accelerated text-to-image generative models that apply the consistency-models framework of Song et al.
A latent space is the vector space a machine learning model maps its inputs into, where each input becomes a point (a latent vector or latent code) and the geometry of the space carries information the raw…
A latent diffusion model (LDM) is a type of diffusion model that runs the denoising diffusion process in a compressed latent space learned by a pretrained autoencoder, rather than directly in pixel space…
Leonardo.AI (stylized as Leonardo.Ai) is a Sydney-based generative AI platform for image generation, video creation, and creative design, founded in December 2022 by CEO JJ Fiasson and five co-founders and…
Luma AI (operating as Luma Labs, Inc.) is a generative AI company headquartered in Palo Alto, California, that builds text-to-video, 3D, and image models, best known for its Dream Machine platform and its Ray…
Luma Dream Machine is a generative AI video and image platform from Luma AI (Luma Labs, Inc.), a San Francisco company, that turns text prompts and still images into short, realistic video clips.
Lumera, short for Light-aware Unified Engine-native Reconstruction and Assembly, is an experimental computer vision benchmark and reference pipeline for turning one RGB image into an editable 3D scene.
Lyria is a family of AI music generation models developed by Google DeepMind, spanning text-to-music synthesis, real-time interactive music performance, and full-length song composition.
Lyria 2 is a high-fidelity, text-to-music generation model built by Google DeepMind that turns text prompts into professional-grade instrumental audio.
Lyria 3.5 is a music generation model from Google DeepMind and the newest member of the Lyria family.
MAGI-2 Preview is a public research release of a unified audio-video generation model developed by Sand.ai.
MAI-Voice-1 is a text-to-speech (speech generation) model developed by Microsoft AI, the consumer artificial-intelligence division of Microsoft led by Mustafa Suleyman.
Magnific AI is an AI-powered image upscaling and enhancement platform founded in November 2023 by Javi Lopez and Emilio Nicolas in Murcia, Spain.
Make-A-Scene is a text-to-image generation model published by Meta AI (then Meta AI Research) in 2022.
Make-A-Video is a text-to-video generation system from Meta AI, announced on September 29, 2022
Marble is a multimodal generative world model developed by World Labs, the spatial intelligence startup co-founded by Stanford computer scientist Fei-Fei Li.
MaskGIT, short for Masked Generative Image Transformer, is an image-synthesis method introduced by Google Research in the 2022 paper "MaskGIT: Masked Generative Image Transformer" by Huiwen Chang, Han Zhang…
Masked Autoregressive (MAR) generation is an image-generation method introduced in the 2024 paper "Autoregressive Image Generation without Vector Quantization" by Tianhong Li, Yonglong Tian, He Li, Mingyang…
Meshy 6 is the sixth major release of Meshy AI's generative 3D generation platform, a hosted generative AI service that turns text prompts or images into textured 3D models (meshes) for games, animation, and…
Midjourney is an artificial intelligence image generation service and independent research lab headquartered in San Francisco, California
Midjourney V7 is the seventh major text-to-image generation model developed by Midjourney Inc. It launched in alpha on April 3, 2025, and became the platform's default model on June 17, 2025.
Minimax loss is a loss function rooted in game theory and decision theory that measures the worst-case performance of a strategy, algorithm, or model.
Movie Gen is a suite of media-generation foundation models from Meta AI, announced on October 4, 2024, that generates high-definition video with synchronized audio from text prompts.
Muse Image is a proprietary image-generation and editing system developed by Meta Superintelligence Labs, a division of Meta.
MuseNet is a deep neural network for symbolic music generation, announced by OpenAI on April 25, 2019.
AI in music is the use of artificial intelligence, especially machine learning and generative AI, to compose, perform, mix, master, transcribe, voice-clone, and reproduce music.
MusicGen is an open-weights text-to-music generation model from Meta AI's Fundamental AI Research (FAIR) team, released on 8 June 2023, that generates roughly 30-second music clips from a text prompt, a…
MusicLM is a text-to-music generation model from Google Research that generates high-fidelity music at 24 kHz from natural language descriptions and keeps that audio consistent over several minutes.
NVIDIA Cosmos is a world foundation model platform developed by NVIDIA for physical AI applications, including autonomous vehicles and robotics.
NVIDIA Cosmos 3 is an open family of "world foundation models" for physical AI that NVIDIA launched on June 1, 2026 at GTC Taipei, held alongside COMPUTEX 2026.
NVIDIA DLSS 5 is an optional real-time generative AI rendering feature developed by Nvidia for video games.
NVIDIA Picasso is a cloud-based generative AI foundry from NVIDIA for building, training, and deploying visual generative models that produce images, video, and 3D content from text prompts.
Nano Banana is the codename, later turned official brand, for Google's native image generation and editing models built into the Gemini ecosystem and developed by Google DeepMind.
Nano Banana 2 is the public nickname for Gemini 3.1 Flash Image, an image generation and editing model released by Google DeepMind on 26 February 2026 .
Nano Banana 2 Lite is a text-to-image generation and image editing model released by Google DeepMind on June 30, 2026.
Nano Banana Pro is a professional grade image generation and editing model from Google DeepMind, released on November 20, 2025.
Natural language generation (NLG) is the subfield of natural language processing and artificial intelligence concerned with building systems that produce understandable text in English or other human…
VALL-E is a zero-shot learning text-to-speech (TTS) system from Microsoft Research that clones a target voice from a 3-second recording and synthesizes new speech in that voice without any per-speaker training.
Odyssey-2 is a family of real-time, interactive generative models built by the AI lab Odyssey.
OpusClip is an AI video-repurposing tool that turns long-form videos into short, social-ready clips, ranking the moments most likely to go viral and adding animated captions, automatic reframing, B-roll, and…
Parti (Pathways Autoregressive Text-to-Image) is a text-to-image generation model from Google Research that produces images from natural-language descriptions by treating the task as a sequence-to-sequence…
Comet is an AI-powered, agentic web browser developed by Perplexity AI. It is built on the Chromium open-source project and ships with an integrated AI assistant that can summarise pages, answer questions…
Perplexity Finance is a dedicated financial research product from Perplexity AI, available at perplexity.ai/finance, that combines real-time market data, company filings, earnings call transcripts, and…
Perplexity Max is the premium subscription tier offered by Perplexity AI, priced at $200 per month or $2,000 per year.
Phind was an AI-powered answer engine built for software developers. The product combined a live web index with fine-tuned large language models to return cited, code-aware answers to programming questions…