Computer Vision

Explore Computer Vision through related topics and the articles other pages reference most.

Most referenced in this topic

Ranked by links from other AI Wiki pages.

Explore articles

Browse subtopics (51)

Articles that also belong to these categories. Counts cover all of Computer Vision.

Showing 1-60 of 206 articles

AI in collectibles

Artificial intelligence in collectibles refers to the use of machine learning and computer vision to grade, authenticate, price, and catalog collectible items such as trading cards, coins, sports memorabilia…

AI Tools & Products

AlexNet

AlexNet is a deep learning convolutional neural network, built by Alex Krizhevsky, Ilya Sutskever, and Geoffrey Hinton at the University of Toronto, that won the ImageNet Large Scale Visual Recognition…

Deep LearningMachine Learning

Augmented reality

Augmented reality (AR) is a class of display and interaction technologies that overlay computer-generated content onto a user's perception of the physical world, keeping the real world visible while adding…

BELEBELE

Belebele is a multiple-choice machine reading comprehension (MRC) AI benchmark that is fully parallel across 122 language variants, meaning the same questions, passages, and answer choices are translated into…

AI Benchmarks

BLINK

BLINK is an AI benchmark that evaluates the core visual perception abilities of multimodal large language models (MLLMs).

AI Benchmarks

Ben Mildenhall

Ben Mildenhall is an American computer scientist known for his work in computer graphics and 3D computer vision

People

BigGAN

BigGAN is a class-conditional generative adversarial network that, when introduced by DeepMind researchers Andrew Brock, Jeff Donahue, and Karen Simonyan in 2018, set a new state of the art for AI image…

Generative AIGoogle DeepMind

Bounding Box

A bounding box is a rectangular region defined by a set of coordinates that encloses an object of interest within an image, video frame, or three-dimensional space.

Machine Learning

CIDEr

CIDEr (Consensus-based Image Description Evaluation) is an automatic evaluation metric for image captioning that scores a machine-generated caption by how closely it matches the consensus of several human…

Machine LearningModel Evaluation

CIFAR-10

CIFAR-10 is a labeled dataset of 60,000 small color images sorted into 10 mutually exclusive object categories, with 6,000 images per class, used as a standard benchmark for image classification.

AI BenchmarksData & Datasets

CLIP Score

CLIP Score (also written CLIPScore or CLIP-S) is a reference-free automatic evaluation metric that measures how well a text caption matches an image, computed as the rescaled cosine similarity of the image and…

AI BenchmarksImage Generation

CloudWalk Technology

CloudWalk Technology Co., Ltd. (Chinese: 云从科技, stock code: 688327.SS) is a Chinese artificial intelligence company, founded in April 2015 by Zhou Xi, that builds facial recognition, computer vision, and…

AI CompaniesChinese AI

Computer-use agent

A computer-use agent (CUA) is a category of AI agent in artificial intelligence that performs tasks by directly operating a general-purpose computer's graphical user interface (GUI) the way a human does, by…

AI AgentsArtificial Intelligence

Convolutional Filter

A convolutional filter (also called a kernel or feature detector) is a small matrix of learnable weights that slides across an input and computes a dot product at each position to produce a feature map.

Deep LearningMachine Learning

Convolutional Layer

A convolutional layer is the core building block of a convolutional neural network (CNN): it slides a small set of learnable filters (also called kernels) across the input, computing a convolution (technically…

Deep LearningMachine Learning

DETR

DETR (DEtection TRansformer) is an end-to-end object detection model that reframes detection as a direct set prediction problem solved with a transformer encoder-decoder and bipartite matching, removing the…

Deep LearningTransformer Models

DINO (computer vision)

DINO (self-DIstillation with NO labels) is a family of self-supervised learning methods for computer vision from Meta AI that trains Vision Transformers (ViTs) on unlabeled images and produces general-purpose…

Deep LearningMachine Learning

DINOv2

DINOv2 is a family of self-supervised Vision Transformer models released by Meta AI Research in April 2023 that produces general-purpose visual features transferring to many downstream tasks without…

AI ModelsMeta AI

DINOv3

DINOv3 is a family of self-supervised computer vision foundation models released by Meta AI in August 2025.

AI ModelsMeta AI

DeepLab

DeepLab is a family of deep convolutional neural network architectures for semantic segmentation, developed by Liang-Chieh Chen and collaborators at UCLA and Google between 2014 and 2018 .

Deep LearningGoogle

DeepSeek-OCR

DeepSeek-OCR is an open-source optical character recognition (OCR) and document-understanding system released by DeepSeek on 20 October 2025 that pioneers a contexts optical compression paradigm: it encodes…

Chinese AIMultimodal AI

Deepfake

A deepfake is synthetic media in which a real person's face, voice, or body is digitally replaced, manipulated, or fabricated using artificial intelligence, most often deep learning techniques such as…

AI EthicsArtificial Intelligence

DeiT

DeiT (Data-efficient Image Transformers) is a family of vision transformer models that proved Vision Transformers can be trained to state-of-the-art image classification accuracy on ImageNet alone

Deep LearningTransformer Models

DenseNet

DenseNet (Densely Connected Convolutional Networks) is a convolutional neural network architecture that connects every layer to every other layer in a feed-forward fashion

Deep LearningNeural Networks

Depth estimation

Depth estimation is the computer vision task of predicting how far each surface in a scene is from the camera, producing a dense per-pixel depth map from one or more images.

Deep Learning

Detectron2

Detectron2 is an open-source software library for object detection and image segmentation, built on PyTorch and developed by Facebook AI Research (FAIR), the research group now part of Meta AI.

Meta AIOpen Source AI

Diffusion model

A diffusion model is a generative model that learns to transform samples from a simple reference distribution into samples resembling a data distribution by reversing a gradual corruption process.

Deep LearningGenerative AI

Donut (Model)

Donut (Document understanding transformer) is an OCR-free visual document understanding model introduced by researchers at NAVER CLOVA in the paper "OCR-free Document Understanding Transformer," first posted…

AI ModelsMultimodal AI

Drone

A drone is an aircraft that flies without a human pilot on board, controlled either remotely by an operator or autonomously by onboard computers using artificial intelligence and sensors.

Robotics

EfficientNet

EfficientNet is a family of convolutional neural network architectures and a model-scaling method that uniformly scales network depth, width, and input resolution with a single compound coefficient, developed…

Deep LearningNeural Networks

Ego-Exo4D

Ego-Exo4D is a large-scale, multimodal, multiview video dataset and benchmark suite for computer vision research on skilled human activity, a central resource in egocentric vision.

Data & DatasetsMeta AI

Ego4D

Ego4D is a large-scale egocentric (first-person) video dataset and benchmark suite for computer vision, assembled by Meta AI (then Facebook AI Research) together with a consortium of 13 universities and labs…

Data & DatasetsMeta AI

EgoSchema

EgoSchema is a diagnostic benchmark for evaluating very long-form video language understanding, introduced by Karttikeya Mangalam, Raiymbek Akshulakov, and Jitendra Malik at UC Berkeley.

AI BenchmarksMultimodal AI

Egocentric Vision

Egocentric vision, also called first-person vision, is the branch of computer vision concerned with video and sensor data recorded from a camera worn on the head or body of the person whose activity is being…

AI Research

European Conference on Computer Vision

The European Conference on Computer Vision (ECCV) is the biennial top-tier academic conference on computer vision, held in even-numbered years at locations across Europe, and is one of the three most…

AI Events