DINOv2
DINOv2 is a family of self-supervised Vision Transformer models released by Meta AI Research in April 2023 that produces general-purpose visual features transferring to many downstream tasks without…
Explore AI Models through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of AI Models.
Showing 1-39 of 39 articles
DINOv2 is a family of self-supervised Vision Transformer models released by Meta AI Research in April 2023 that produces general-purpose visual features transferring to many downstream tasks without…
DINOv3 is a family of self-supervised computer vision foundation models released by Meta AI in August 2025.
Donut (Document understanding transformer) is an OCR-free visual document understanding model introduced by researchers at NAVER CLOVA in the paper "OCR-free Document Understanding Transformer," first posted…
Florence-2 is a vision foundation model developed by Microsoft Research that handles a wide range of computer vision and vision-language tasks through a single unified
GAIA-2 (Generative AI for Autonomy 2) is a controllable, multi-camera generative world model for autonomous driving, announced by the British self-driving company Wayve on 26 March 2025.
GAIA-3 is a 15-billion-parameter generative world model for autonomous driving released by Wayve on 2 December 2025, the third generation in the company's GAIA family.
GAIA-4 is a multimodal generative world model for closed-loop autonomous driving simulation, announced by the British self-driving company Wayve on 3 August 2026 as the latest generation of its GAIA family.
GoogLeNet, also known as Inception v1, is a 22-layer deep convolutional neural network introduced by Google researchers in 2014 that won the classification task of the ImageNet Large Scale Visual Recognition…
Grok Imagine is a generative media product from xAI, the company founded by Elon Musk.
Ideogram 3.0 is a text-to-image generation model released by Ideogram on March 26, 2025.
Image classification models are machine learning systems that assign one or more category labels to a whole input image, the task that drove the modern wave of deep learning in computer vision.
Image-to-image models (often shortened to img2img) are machine learning systems that take an input image and output a transformed version of it.
InternVideo is a family of general-purpose video foundation models developed by OpenGVLab at the Shanghai Artificial Intelligence Laboratory in collaboration with Nanjing University and the Shenzhen Institutes…
Luma Dream Machine is a generative AI video and image platform from Luma AI (Luma Labs, Inc.), a San Francisco company, that turns text prompts and still images into short, realistic video clips.
Mistral OCR 3 is a document-understanding and optical character recognition model from Mistral AI, released in mid-December 2025 as the third generation of the company's OCR product line.
Mistral OCR 4 is a proprietary document extraction model and service developed by Mistral AI. Released on June 23, 2026, it converts documents and page images into Markdown and structured layout data.
Nano Banana is the codename, later turned official brand, for Google's native image generation and editing models built into the Gemini ecosystem and developed by Google DeepMind.
Nougat (Neural Optical Understanding for Academic Documents) is a document-understanding model from Meta AI that converts the rendered image of a document page into structured markup text.
Pika 2.5 is a generative video model developed by Pika Labs, the San Francisco based AI video startup co-founded in April 2023 by Stanford AI Lab dropouts Demi Guo and Chenlin Meng.
Robostral Navigate is an 8-billion-parameter vision language model developed by Mistral AI for instruction-following robot navigation.
Runway Act-Two is a generative motion capture and character animation model developed by Runway, publicly introduced on July 15, 2025.
Runway Aleph is an in-context AI video editing model from Runway that transforms and edits an existing video clip from a plain-text instruction, performing a wide range of tasks in a single model: adding…
SAM 2 (Segment Anything Model 2) is a promptable visual segmentation model for both images and video developed by Meta AI and released on 29 July 2024.
Sapiens is a family of human-centric computer vision foundation models developed by Meta (Reality Labs), introduced in 2024 and presented as an oral paper at the European Conference on Computer Vision (ECCV)…
Seedance is the family of foundation video generation models built by the Seed team at ByteDance, the Chinese internet company that owns TikTok and Douyin.
Seedance 2.5 is a proprietary AI video generation model released by ByteDance on July 31, 2026. It jointly generates audio and video from combinations of text, image, video, and audio inputs.
Seedream is a series of text-to-image and image-editing foundation models built by the Seed research team at ByteDance, the company behind TikTok and Douyin.
Segment Anything Model (SAM) is a promptable image segmentation foundation model released by Meta AI on April 5, 2023 that lets users "cut out" any object in an image with a single click, box, or mask prompt…
Synthesia 3.0 is a major release of the AI video generation platform from Synthesia, the London-based company co-founded in 2017 by Victor Riparbelli, Steffen Tjerrild, Lourdes Agapito and Matthias Niessner.
Unconditional image generation models are generative neural networks that learn the marginal distribution p(x) of a set of training images and produce new samples from that learned distribution, with no extra…
V-JEPA (Video Joint Embedding Predictive Architecture) is a self-supervised video model from Meta AI that learns by predicting masked regions of a video in an abstract latent representation space rather than…
V-JEPA 2 (Video Joint Embedding Predictive Architecture 2) is an open-source video world model released by Meta AI on June 11, 2025 that learns to understand, predict, and plan in the physical world by…
VGGNet is a deep convolutional neural network architecture, introduced in 2014 by Karen Simonyan and Andrew Zisserman of the Visual Geometry Group at the University of Oxford, that classifies images using a…
Video classification models are machine learning systems that assign one or more category labels to a video clip, typically describing the human action depicted.
Visual question answering models are AI systems that take an image and a natural language question about that image and return a natural language answer.
Wan 2.1-VACE (also written Wan2.1-VACE) is an open-weights video creation and editing model released by Alibaba's Tongyi Lab on May 14, 2025 .
Wan 2.5 is a natively multimodal AI video generation model developed by Alibaba Cloud's Tongyi Lab and previewed at the company's Apsara 2025 conference in Hangzhou on September 24, 2025.
YOLOv8 is a family of one-stage computer vision models released by Ultralytics on January 10, 2023.
Zero-shot image classification models are vision systems that assign images to categories the model has never encountered as labeled training examples.