Black Forest Labs
Black Forest Labs (BFL) is a German-American artificial intelligence company founded in 2024 by Robin Rombach, Andreas Blattmann, Patrick Esser, and Dominik Lorenz
Explore Diffusion Models through related topics and the articles other pages reference most.
Ranked by links from other AI Wiki pages.
Articles that also belong to these categories. Counts cover all of Diffusion Models.
Showing 1-43 of 43 articles
Black Forest Labs (BFL) is a German-American artificial intelligence company founded in 2024 by Robin Rombach, Andreas Blattmann, Patrick Esser, and Dominik Lorenz
Consistency models are a family of generative models, introduced by Yang Song, Prafulla Dhariwal, Mark Chen, and Ilya Sutskever at OpenAI in March 2023
DALL-E is a family of text-to-image systems developed by OpenAI that generates images from natural-language descriptions.
Denoising Diffusion Implicit Models (DDIM) are a class of iterative generative models, introduced by Jiaming Song, Chenlin Meng, and Stefano Ermon of Stanford University in October 2020
Diffusion language models (DLMs, sometimes written dLLMs at frontier scale) are text generators that synthesize a sequence by iteratively denoising or unmasking many tokens in parallel
A Diffusion Transformer (DiT) is a transformer-based neural network backbone for diffusion models that replaces the U-Net with a Vision Transformer operating on patches of an image latent.
Diffusion policy is a robot imitation-learning method, introduced in 2023 by Cheng Chi, Shuran Song, and collaborators at Columbia University, the Toyota Research Institute, and MIT
DiffusionGemma is an experimental open-weight large language model developed by Google DeepMind for multimodal text generation through discrete diffusion.
A discrete diffusion language model is a class of generative model for text that produces tokens by iteratively denoising a corrupted sequence, rather than by predicting one token at a time from left to right.
EDM is the common shorthand for the paper "Elucidating the Design Space of Diffusion-Based Generative Models" by Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine of NVIDIA, presented at NeurIPS 2022.
FLUX.1 is a family of text-to-image generation models developed by Black Forest Labs, released on August 1, 2024.
FLUX.2 is the second-generation image generation and editing model family developed by Black Forest Labs, released on November 25, 2025.
Flow Matching is a simulation-free training framework for generative models that fits a time-dependent velocity field to transport samples from a source distribution (typically a standard Gaussian) to a data…
Flux is a family of text-to-image generative models developed by Black Forest Labs (BFL), the German-American startup founded by the original creators of Stable Diffusion.
GLIDE (Guided Language to Image Diffusion for Generation and Editing) is a text-conditional diffusion model for text-to-image synthesis and editing released by OpenAI in December 2021.
Gemini Diffusion is an experimental text generation model from Google DeepMind that produces text and code using a diffusion process rather than the autoregressive
GenCast is a probabilistic, ensemble-based machine-learning weather forecasting model developed by Google DeepMind that uses a diffusion model to generate large ensembles of possible 15-day weather…
Imagen is a family of text-to-image diffusion models developed by Google, first introduced in May 2022 and as of 2026 in its fourth generation (Imagen 4).
Imagen 2 is the second generation of Google's text-to-image diffusion model, developed by Google DeepMind and first announced for developers and enterprises on December 13, 2023.
Inception Labs (often referred to simply as Inception) is a Palo Alto, California-based artificial intelligence startup that commercializes diffusion language models (dLLMs) for text and code generation.
LLaDA (Large Language Diffusion with mAsking) is a family of non-autoregressive large language models that generate text by iteratively denoising a sequence of mask tokens rather than predicting tokens left to…
LTX-Video is an open-source, transformer-based latent video diffusion model developed by the Israeli company Lightricks and first released to the public in November 2024.
Latent Consistency Models (LCMs) are a family of accelerated text-to-image generative models that apply the consistency-models framework of Song et al.
Lumiere is a text-to-video diffusion model developed by Google Research in collaboration with researchers from the Weizmann Institute of Science, Tel Aviv University, and the Technion
MMDiT (Multimodal Diffusion Transformer, sometimes written MM-DiT) is a transformer architecture for text-conditioned image generation that gives image tokens and text tokens their own separate weights but…
Mercury is a family of commercial-scale diffusion-based large language models developed by Inception Labs, a Palo Alto startup co-founded in 2024 by Stanford professor Stefano Ermon together with his former…
Midjourney is an artificial intelligence image generation service and independent research lab headquartered in San Francisco, California
Mochi 1 is an open-weights text-to-video diffusion model released by Genmo Inc. on October 22, 2024 under the Apache 2.0 license.
Nemotron-Labs-TwoTower is an open-weight diffusion language model released by NVIDIA in mid-2026.
Open-Sora is an open-source text-to-video diffusion project initiated in March 2024 by Singapore-based startup HPC-AI Tech (the team behind the Colossal-AI distributed training framework) as a public attempt…
RFdiffusion (short for RoseTTAFold diffusion) is a deep-learning system for de novo protein design developed at the University of Washington Institute for Protein Design (IPD) in David Baker's lab.
Rectified Flow is a generative modeling framework that learns a transport ordinary differential equation between two probability distributions by regressing a velocity field along straight-line interpolations…
Runwayml/stable-diffusion-v1-5 is the Hugging Face repository name of the Stable Diffusion v1.5 checkpoint, a text-to-image latent diffusion model published on October 20
SDXL, short for Stable Diffusion XL, is an open-weights latent text-to-image diffusion model released by Stability AI on 26 July 2023, built around a 2.6 billion parameter U-Net backbone, two text encoders…
Score Entropy Discrete Diffusion (SEDD) is a discrete diffusion model for language and other discrete data introduced by Aaron Lou, Chenlin Meng, and Stefano Ermon at Stanford University in the paper Discrete…
Score matching is a method for fitting a probabilistic model by matching the gradient of its log-density, the so-called score function $$\nabla_x \log p(x)$$, to the same gradient of the data distribution
Sora was a text-to-video generation model developed by OpenAI. OpenAI revealed the research model on February 15, 2024 and granted early access to red teamers, visual artists, designers, and filmmakers.
Stable Diffusion is a family of generative image models that can synthesize and edit images from text and other conditions.
Stable Diffusion 3 (SD3) is a family of text-to-image diffusion models developed by Stability AI, first announced as an early preview on February 22, 2024, and built on a new architecture called the Multimodal…
Stable Diffusion 3.5 (SD 3.5) is a family of open-weights text-to-image diffusion models released by Stability AI on October 22, 2024, comprising three variants: Stable Diffusion 3.5 Large (8.1 billion…
Stable Video Diffusion (SVD) is a latent video diffusion model released by Stability AI on 21 November 2023
Text-to-video (often abbreviated T2V) is the generative AI capability of producing video clips, with or without sound, directly from a written prompt.
Würstchen is an efficient three-stage cascaded latent diffusion architecture for text-to-image synthesis introduced by Pablo Pernias, Dominic Rampas, Mats L. Richter, Christopher J. Pal