Megvii
Megvii Technology Limited (Chinese: 旷视科技), also known by the name of its flagship product Face++, is a Beijing-based artificial intelligence company specializing in computer vision, deep learning, and Internet…
Explore Computer Vision through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of Computer Vision.
Showing 121-180 of 206 articles
Megvii Technology Limited (Chinese: 旷视科技), also known by the name of its flagship product Face++, is a Beijing-based artificial intelligence company specializing in computer vision, deep learning, and Internet…
Mistral OCR 3 is a document-understanding and optical character recognition model from Mistral AI, released in mid-December 2025 as the third generation of the company's OCR product line.
Mistral OCR 4 is a proprietary document extraction model and service developed by Mistral AI. Released on June 23, 2026, it converts documents and page images into Markdown and structured layout data.
MobileNet is a family of efficient convolutional neural network (CNN) architectures developed by Google for mobile and edge AI applications.
Mobileye Global Inc. (Nasdaq: MBLY) is an Israeli technology company headquartered in Jerusalem that develops vision-based advanced driver-assistance systems (ADAS) and autonomous driving technology
A multimodal model is a machine learning model, or a model-based system, that processes, relates, or produces information across more than one kind of data. Each kind is called a modality.
NVIDIA DeepStream is a GPU-accelerated software development kit for building real-time streaming analytics pipelines, principally for video.
Nano Banana is the codename, later turned official brand, for Google's native image generation and editing models built into the Gemini ecosystem and developed by Google DeepMind.
Neural Radiance Fields (NeRF) is a method for synthesizing photorealistic novel views of a 3D scene by encoding the scene as a continuous 5D function (3D position plus 2D viewing direction) inside a single…
Niantic Spatial, Inc. is an American geospatial artificial intelligence company that builds maps and 3D reconstructions of the physical world for machines, including humanoid robots, delivery robots, augmented…
Nougat (Neural Optical Understanding for Academic Documents) is a document-understanding model from Meta AI that converts the rendered image of a document page into structured markup text.
OCR Models are artificial intelligence (AI) systems that convert images of typed, handwritten, or printed text into machine-readable digital text through Optical Character Recognition (OCR).
Object detection is a computer vision task that finds instances of interest in an image and assigns each one a category.
Open Images is a large annotated image dataset published by Google for computer vision research.
OpenGVLab is the open-source organization and project hub for general-vision and multimodal foundation models run by the General Vision group at Shanghai AI Laboratory.
OpenPose is an open-source library for real-time multi-person 2D pose estimation that detects body, foot, hand, and facial keypoints in images and video.
PASCAL VOC (Pattern Analysis, Statistical Modelling and Computational Learning Visual Object Classes) is a long-running benchmark dataset and annual challenge for object recognition, object detection…
Panoptic segmentation is a computer vision task that assigns every pixel in an image a semantic class label and, for countable objects, an instance identity.
Perception Encoder (PE) is a family of vision and vision-language encoders from Meta AI's Fundamental AI Research (FAIR) group, released in April 2025
Artificial intelligence in photography covers a wide span of techniques, from the computational pipelines baked into modern smartphones to the generative editing tools now built into Photoshop and Lightroom…
Photoroom is an AI-powered photo editing platform headquartered in Paris, France, specializing in background removal, product photography, and generative image editing.
Pika 2.5 is a generative video model developed by Pika Labs, the San Francisco based AI video startup co-founded in April 2023 by Stanford AI Lab dropouts Demi Guo and Chenlin Meng.
Pooling is a downsampling operation in neural networks that aggregates each local region of a feature map into a single summary value
Pose estimation is the computer vision task of detecting and localizing the keypoints (also called landmarks or joints) of a human body, hand, face, animal, or rigid object in images and video, then connecting…
Pre-training is a stage of machine learning in which a model learns parameters from a source dataset or source objective before those parameters are reused or adapted for a target use.
Project Aria is an egocentric data-collection research program run by Meta's Reality Labs Research. It was announced on September 16, 2020 .
R-CNN (short for Regions with CNN features) is a two-stage object detection method that generates about 2,000 candidate region proposals per image with Selective Search, warps each region and runs a…
Rerun is an open-source multimodal data stack and visualization system built for robotics, computer vision, and other forms of physical AI.
ResNet, short for residual network, is a family of deep convolutional neural networks introduced by Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun at CVPR 2016.
Robostral Navigate is an 8-billion-parameter vision language model developed by Mistral AI for instruction-following robot navigation.
Robot manipulation is the ability of a robotic system to physically interact with objects in its environment through grasping, pushing, pulling, inserting, placing, and other contact-based actions.
Robot navigation is the ability of a mobile robot to determine its own position in its environment, plan a path toward a goal, and execute that path while avoiding obstacles.
Runway Act-Two is a generative motion capture and character animation model developed by Runway, publicly introduced on July 15, 2025.
Runway Aleph is an in-context AI video editing model from Runway that transforms and edits an existing video clip from a plain-text instruction, performing a wide range of tasks in a single model: adding…
SAM 2 (Segment Anything Model 2) is a promptable visual segmentation model for both images and video developed by Meta AI and released on 29 July 2024.
SSD (Single Shot MultiBox Detector) is a single stage object detection model that predicts bounding boxes and per class confidence scores in one forward pass through a convolutional network
A saliency map is an explainable AI visualization that highlights which parts of an input, most often the individual pixels of an image, most influenced a deep learning model's prediction.
Sand.ai is an artificial-intelligence company that develops video-generation models, research software, and commercial creation tools.
Sapiens is a family of human-centric computer vision foundation models developed by Meta (Reality Labs), introduced in 2024 and presented as an oral paper at the European Conference on Computer Vision (ECCV)…
Seedance is the family of foundation video generation models built by the Seed team at ByteDance, the Chinese internet company that owns TikTok and Douyin.
Seedance 2.5 is a proprietary AI video generation model released by ByteDance on July 31, 2026. It jointly generates audio and video from combinations of text, image, video, and audio inputs.
Seedream is a series of text-to-image and image-editing foundation models built by the Seed research team at ByteDance, the company behind TikTok and Douyin.
Segment Anything Model (SAM) is a promptable image segmentation foundation model released by Meta AI on April 5, 2023 that lets users "cut out" any object in an image with a single click, box, or mask prompt…
Semantic segmentation is a computer vision task that assigns a category label to every single pixel in an image, producing a dense map in which each pixel carries the identity of the object class it belongs to.
SenseTime (Chinese: 商汤科技; pinyin: Shāngtāng Kējì) is a partly state-owned, publicly traded artificial intelligence company headquartered in Hong Kong that grew into China's largest AI software company by…
SigLIP (Sigmoid Loss for Language-Image Pre-training) is a family of vision-language encoders developed by researchers at Google DeepMind that pre-trains image and text encoders by treating each image-text…
SimCLR (Simple Framework for Contrastive Learning of Visual Representations) is a self-supervised learning method for computer vision in which a network is trained to recognise that two differently augmented…
Simultaneous Localization and Mapping (SLAM) is the computational problem of building a map of an unknown environment while at the same time estimating the position of a sensor, vehicle, or agent moving…
Size invariance, also called scale invariance, is the property of a model, feature, or algorithm that produces the same output (or class label) regardless of the size at which an object appears in the input.
Skydio is an American drone manufacturer based in San Mateo, California, that builds autonomous flying robots guided by onboard computer vision rather than manual piloting.
Spatial pooling is a downsampling operation in convolutional neural networks (CNNs) that replaces a local region of a feature map with a single summary statistic, such as the maximum or the average of the…
Spatial intelligence is the ability of an AI system to perceive, understand, reason about, generate, and interact with three-dimensional space rather than just text or two-dimensional pixels.
Stride is the step size by which a filter (or pooling window) moves across the input in a convolutional neural network (CNN): a stride of 1 shifts the filter one position at a time and visits every location…
StyleGAN is a family of style-based generative adversarial network (GAN) architectures developed by NVIDIA Research for high-quality unconditional image synthesis
The Swin Transformer (Shifted Window Transformer) is a hierarchical vision transformer architecture that computes self-attention within local, non-overlapping windows and introduces a shifted window…
Synthesia 3.0 is a major release of the AI video generation platform from Synthesia, the London-based company co-founded in 2017 by Victor Riparbelli, Steffen Tjerrild, Lourdes Agapito and Matthias Niessner.
T2I-CompBench is an AI benchmark for evaluating compositional text-to-image generation, introduced in the paper "T2I-CompBench: A Comprehensive Benchmark for Open-world Compositional Text-to-image Generation"…
Topaz Labs is a software company headquartered in Dallas, Texas, that develops AI-powered tools for photo and video enhancement, best known for Gigapixel, Photo AI, Video AI, and its Bloom and Project…
Translational invariance (also called translation invariance or shift invariance) is the property of a function, system, or machine learning model whose output does not change when its input is translated…
U-Net is a convolutional neural network architecture designed for biomedical image segmentation.