Agentic video understanding
Agentic video understanding is an approach to video analysis in which a reasoning model chooses which parts of a video to inspect for a particular question.
Explore Multimodal AI through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of Multimodal AI.
Showing 1-7 of 7 articles
Agentic video understanding is an approach to video analysis in which a reasoning model chooses which parts of a video to inspect for a particular question.
CLIP, short for Contrastive Language-Image Pre-training, is a family of neural networks developed by OpenAI researchers to learn a shared representation of images and natural-language descriptions.
Claude Sonnet 4.5 is a multimodal large language model (LLM) developed by Anthropic and released on September 29, 2025, which Anthropic described at launch as "the best coding model in the world." It is a…
A computer-use agent (CUA) is a category of AI agent in artificial intelligence that performs tasks by directly operating a general-purpose computer's graphical user interface (GUI) the way a human does, by…
Image-to-text models are machine learning systems that take an image as input and produce natural language text as output.
MMMU (Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark) is a multimodal AI benchmark of 11,550 college-level questions that pairs text with images to test expert knowledge and…
data2vec is a self-supervised learning framework from Meta AI (then Facebook AI Research) that applies the same training method to three different input types: speech, computer vision, and text.