AI Models

Explore AI Models through related topics and the articles other pages reference most.

Explore articles

Reset filters
Browse subtopics: Computer Vision

Articles that also belong to these categories. Counts cover all of AI Models.

Showing 1-39 of 39 articles

DINOv2

DINOv2 is a family of self-supervised Vision Transformer models released by Meta AI Research in April 2023 that produces general-purpose visual features transferring to many downstream tasks without…

Computer VisionMeta AI

Donut (Model)

Donut (Document understanding transformer) is an OCR-free visual document understanding model introduced by researchers at NAVER CLOVA in the paper "OCR-free Document Understanding Transformer," first posted…

Computer VisionMultimodal AI

Florence-2

Florence-2 is a vision foundation model developed by Microsoft Research that handles a wide range of computer vision and vision-language tasks through a single unified

Computer VisionMicrosoft

GAIA-4 (Wayve)

GAIA-4 is a multimodal generative world model for closed-loop autonomous driving simulation, announced by the British self-driving company Wayve on 3 August 2026 as the latest generation of its GAIA family.

Autonomous VehiclesComputer Vision

GoogLeNet (Inception v1)

GoogLeNet, also known as Inception v1, is a 22-layer deep convolutional neural network introduced by Google researchers in 2014 that won the classification task of the ImageNet Large Scale Visual Recognition…

AI HistoryComputer Vision

Image Classification Models

Image classification models are machine learning systems that assign one or more category labels to a whole input image, the task that drove the modern wave of deep learning in computer vision.

Computer Vision

Image-to-Image Models

Image-to-image models (often shortened to img2img) are machine learning systems that take an input image and output a transformed version of it.

Computer Vision

InternVideo

InternVideo is a family of general-purpose video foundation models developed by OpenGVLab at the Shanghai Artificial Intelligence Laboratory in collaboration with Nanjing University and the Shenzhen Institutes…

Chinese AIComputer Vision

Luma Dream Machine

Luma Dream Machine is a generative AI video and image platform from Luma AI (Luma Labs, Inc.), a San Francisco company, that turns text prompts and still images into short, realistic video clips.

Computer VisionGenerative AI

Mistral OCR 3

Mistral OCR 3 is a document-understanding and optical character recognition model from Mistral AI, released in mid-December 2025 as the third generation of the company's OCR product line.

AI CompaniesComputer Vision

Mistral OCR 4

Mistral OCR 4 is a proprietary document extraction model and service developed by Mistral AI. Released on June 23, 2026, it converts documents and page images into Markdown and structured layout data.

AI CompaniesComputer Vision

Nano Banana

Nano Banana is the codename, later turned official brand, for Google's native image generation and editing models built into the Gemini ecosystem and developed by Google DeepMind.

Computer VisionGenerative AI

Nougat (model)

Nougat (Neural Optical Understanding for Academic Documents) is a document-understanding model from Meta AI that converts the rendered image of a document page into structured markup text.

Computer VisionMeta AI

Pika 2.5

Pika 2.5 is a generative video model developed by Pika Labs, the San Francisco based AI video startup co-founded in April 2023 by Stanford AI Lab dropouts Demi Guo and Chenlin Meng.

Computer VisionGenerative AI

Runway Aleph

Runway Aleph is an in-context AI video editing model from Runway that transforms and edits an existing video clip from a plain-text instruction, performing a wide range of tasks in a single model: adding…

Computer VisionGenerative AI

SAM 2

SAM 2 (Segment Anything Model 2) is a promptable visual segmentation model for both images and video developed by Meta AI and released on 29 July 2024.

Computer VisionMeta AI

Sapiens (computer vision)

Sapiens is a family of human-centric computer vision foundation models developed by Meta (Reality Labs), introduced in 2024 and presented as an oral paper at the European Conference on Computer Vision (ECCV)…

Computer VisionMeta AI

Seedance

Seedance is the family of foundation video generation models built by the Seed team at ByteDance, the Chinese internet company that owns TikTok and Douyin.

Chinese AIComputer Vision

Seedance 2.5

Seedance 2.5 is a proprietary AI video generation model released by ByteDance on July 31, 2026. It jointly generates audio and video from combinations of text, image, video, and audio inputs.

Chinese AIComputer Vision

Seedream

Seedream is a series of text-to-image and image-editing foundation models built by the Seed research team at ByteDance, the company behind TikTok and Douyin.

Chinese AIComputer Vision

Synthesia 3.0

Synthesia 3.0 is a major release of the AI video generation platform from Synthesia, the London-based company co-founded in 2017 by Victor Riparbelli, Steffen Tjerrild, Lourdes Agapito and Matthias Niessner.

Computer VisionGenerative AI

Unconditional Image Generation Models

Unconditional image generation models are generative neural networks that learn the marginal distribution p(x) of a set of training images and produce new samples from that learned distribution, with no extra…

Computer Vision

V-JEPA

V-JEPA (Video Joint Embedding Predictive Architecture) is a self-supervised video model from Meta AI that learns by predicting masked regions of a video in an abstract latent representation space rather than…

Computer VisionMachine Learning

V-JEPA 2

V-JEPA 2 (Video Joint Embedding Predictive Architecture 2) is an open-source video world model released by Meta AI on June 11, 2025 that learns to understand, predict, and plan in the physical world by…

Computer VisionOpen Source AI

VGGNet

VGGNet is a deep convolutional neural network architecture, introduced in 2014 by Karen Simonyan and Andrew Zisserman of the Visual Geometry Group at the University of Oxford, that classifies images using a…

AI HistoryComputer Vision

Wan 2.5

Wan 2.5 is a natively multimodal AI video generation model developed by Alibaba Cloud's Tongyi Lab and previewed at the company's Apsara 2025 conference in Hangzhou on September 24, 2025.

Chinese AIComputer Vision