Computer Vision

Explore Computer Vision through related topics and the articles other pages reference most.

Explore articles

Browse subtopics (51)

Articles that also belong to these categories. Counts cover all of Computer Vision.

Showing 181-206 of 206 articles

Unconditional Image Generation Models

Unconditional image generation models are generative neural networks that learn the marginal distribution p(x) of a set of training images and produce new samples from that learned distribution, with no extra…

AI Models

V-JEPA

V-JEPA (Video Joint Embedding Predictive Architecture) is a self-supervised video model from Meta AI that learns by predicting masked regions of a video in an abstract latent representation space rather than…

AI ModelsMachine Learning

V-JEPA 2

V-JEPA 2 (Video Joint Embedding Predictive Architecture 2) is an open-source video world model released by Meta AI on June 11, 2025 that learns to understand, predict, and plan in the physical world by…

AI ModelsOpen Source AI

VGG

VGG (also called VGGNet) is a deep convolutional neural network architecture, introduced in 2014 by Karen Simonyan and Andrew Zisserman of the Visual Geometry Group at the University of Oxford

Deep LearningNeural Networks

VGGNet

VGGNet is a deep convolutional neural network architecture, introduced in 2014 by Karen Simonyan and Andrew Zisserman of the Visual Geometry Group at the University of Oxford, that classifies images using a…

AI HistoryAI Models

Video Classification Models

Video classification models are machine learning systems that assign one or more category labels to a video clip, typically describing the human action depicted.

AI Models

Video-MME

Video-MME (Video Multi-Modal Evaluation) is a benchmark for testing how well multimodal large language models (MLLMs) understand video, built from 900 manually selected videos totaling 254 hours and 2,700…

AI BenchmarksMultimodal AI

Vision Transformer

The Vision Transformer (ViT) is a deep learning architecture that represents an image as a sequence of fixed-size patches and processes that sequence with a Transformer encoder.

Model Architecture

Vision-based tactile sensor

A vision-based tactile sensor, also called an optical or camera-based tactile sensor, is a category of tactile sensing that measures touch by watching it rather than by reading a strain gauge, a capacitor, or…

Robotics

WISE

WISE (World Knowledge-Informed Semantic Evaluation) is an AI benchmark that tests whether a text-to-image model actually possesses and correctly applies real-world knowledge when it draws a scene

AI Benchmarks

Wan 2.1-VACE

Wan 2.1-VACE (also written Wan2.1-VACE) is an open-weights video creation and editing model released by Alibaba's Tongyi Lab on May 14, 2025 .

AI ModelsChinese AI

Wan 2.5

Wan 2.5 is a natively multimodal AI video generation model developed by Alibaba Cloud's Tongyi Lab and previewed at the company's Apsara 2025 conference in Hangzhou on September 24, 2025.

AI ModelsChinese AI

Wayve

Wayve is a British artificial intelligence company, founded in Cambridge in 2017 by Alex Kendall and Amar Shah, that develops end to end "embodied AI" software for autonomous driving and robotics, and in May…

AI CompaniesAutonomous Vehicles

World Labs

World Labs is an American spatial-intelligence company, headquartered in San Francisco, that builds "Large World Models" (LWMs), generative AI systems that perceive, generate, reason about and interact with…

AI CompaniesGenerative AI

YOLO (object detection)

YOLO (You Only Look Once) is a family of object detection models that treat detection as a single regression problem, predicting bounding boxes and class probabilities directly from full images in one forward…

Deep LearningNeural Networks

Yitu Technology

Yitu Technology (Chinese: 依图科技; pinyin: Yītú Kējì) is a Chinese artificial intelligence company headquartered in Shanghai that was founded in 2012 by Zhu Long (朱珑) and Lin Chenxi (林晨曦)

AI CompaniesChinese AI

ZeroBench

ZeroBench is a visual reasoning benchmark built to be effectively impossible for current frontier large multimodal models, which score 0.0% on its main questions.

AI BenchmarksMultimodal AI

pix2pix

pix2pix is a supervised image-to-image translation method that learns a mapping between two visual domains from aligned input-output image pairs.

Generative AIImage Generation