AI Benchmarks

Explore AI Benchmarks through related topics and the articles other pages reference most.

Explore articles

Reset filters
Browse subtopics: Computer Vision

Articles that also belong to these categories. Counts cover all of AI Benchmarks.

Showing 1-19 of 19 articles

BELEBELE

Belebele is a multiple-choice machine reading comprehension (MRC) AI benchmark that is fully parallel across 122 language variants, meaning the same questions, passages, and answer choices are translated into…

Computer Vision

BLINK

BLINK is an AI benchmark that evaluates the core visual perception abilities of multimodal large language models (MLLMs).

Computer Vision

CIFAR-10

CIFAR-10 is a labeled dataset of 60,000 small color images sorted into 10 mutually exclusive object categories, with 6,000 images per class, used as a standard benchmark for image classification.

Computer VisionData & Datasets

CLIP Score

CLIP Score (also written CLIPScore or CLIP-S) is a reference-free automatic evaluation metric that measures how well a text caption matches an image, computed as the rescaled cosine similarity of the image and…

Computer VisionImage Generation

EgoSchema

EgoSchema is a diagnostic benchmark for evaluating very long-form video language understanding, introduced by Karttikeya Mangalam, Raiymbek Akshulakov, and Jitendra Malik at UC Berkeley.

Computer VisionMultimodal AI

GenAI-Bench

GenAI-Bench is an AI benchmark for evaluating compositional text-to-image and text-to-video generation, introduced in 2024 by researchers from Carnegie Mellon University and Meta AI .

Computer Vision

MM-Vet

MM-Vet is an AI benchmark for evaluating large multimodal models (LMMs, also called multimodal large language models or MLLMs) on tasks that require combining several core vision-language skills at once.

Computer Vision

MMMU-Pro

MMMU-Pro is a rigorous benchmark for evaluating multimodal AI systems on college-level, expert questions that genuinely require seeing an image, built as a harder and more robust version of the original MMMU…

Computer VisionLarge Language Models

PASCAL VOC

PASCAL VOC (Pattern Analysis, Statistical Modelling and Computational Learning Visual Object Classes) is a long-running benchmark dataset and annual challenge for object recognition, object detection…

Computer VisionData & Datasets

T2I-CompBench

T2I-CompBench is an AI benchmark for evaluating compositional text-to-image generation, introduced in the paper "T2I-CompBench: A Comprehensive Benchmark for Open-world Compositional Text-to-image Generation"…

Computer Vision

Video-MME

Video-MME (Video Multi-Modal Evaluation) is a benchmark for testing how well multimodal large language models (MLLMs) understand video, built from 900 manually selected videos totaling 254 hours and 2,700…

Computer VisionMultimodal AI

WISE

WISE (World Knowledge-Informed Semantic Evaluation) is an AI benchmark that tests whether a text-to-image model actually possesses and correctly applies real-world knowledge when it draws a scene

Computer Vision

ZeroBench

ZeroBench is a visual reasoning benchmark built to be effectively impossible for current frontier large multimodal models, which score 0.0% on its main questions.

Computer VisionMultimodal AI