Multimodal AI

Explore Multimodal AI through related topics and the articles other pages reference most.

Explore articles

Browse subtopics (42)

Articles that also belong to these categories. Counts cover all of Multimodal AI.

Showing 61-118 of 118 articles

Llama 3.2

Llama 3.2 is a family of four open-weight large language models released by Meta on September 25, 2024, comprising lightweight 1 billion and 3 billion parameter text-only models for on-device AI and the 11…

AI ModelsLarge Language Models

MM-BrowseComp

MM-BrowseComp is a benchmark for evaluating multimodal web-browsing AI agents, introduced in August 2025 by researchers from ByteDance, Nanjing University, M-A-P, the Institute of Automation of the Chinese…

AI AgentsAI Benchmarks

MMMU

MMMU (Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark) is a multimodal AI benchmark of 11,550 college-level questions that pairs text with images to test expert knowledge and…

AI BenchmarksMachine Learning

MMMU-Pro

MMMU-Pro is a rigorous benchmark for evaluating multimodal AI systems on college-level, expert questions that genuinely require seeing an image, built as a harder and more robust version of the original MMMU…

AI BenchmarksComputer Vision

MMStar

MMStar (Multi-modal Star) is a vision-language model evaluation benchmark consisting of 1,500 multimodal samples that were filtered from six pre-existing benchmarks to ensure both visual dependency (questions…

AI Benchmarks

MathVista

MathVista is a benchmark for evaluating the mathematical reasoning capabilities of foundation models in visual contexts.

AI Benchmarks

MetaCLIP

MetaCLIP (Metadata-Curated Language-Image Pre-training) is a data curation recipe and a family of vision-language models from Meta AI, introduced in the 2023 paper "Demystifying CLIP Data" by Hu Xu, Saining…

Data & DatasetsMeta AI

MiniCPM-V

MiniCPM-V is a family of open-weights multimodal large language models published by OpenBMB, the shared open-source brand of Tsinghua University's Natural Language Processing lab (THUNLP) and the Beijing…

Chinese AIOpen Source AI

Molmo

Molmo is a family of open-weight, open-data vision-language models (VLMs) released by the Allen Institute for AI (Ai2) on 25 September 2024.

AI ModelsOpen Source AI

Muse Glimmer

Muse Glimmer is an open-weight text-and-image model developed by Meta AI for local agent and coding workloads. Meta released the model on August 10, 2026 under the identifier meta-models/Muse-Glimmer-30B.

AI AgentsAI Models

Muse Spark

Muse Spark is a proprietary multimodal reasoning model developed by Meta Superintelligence Labs (MSL), the artificial intelligence division Meta reorganized in 2025.

Meta AIReasoning Models

NVIDIA Cosmos Reason

The original NVIDIA Cosmos Reason release is an open, customizable, 7-billion-parameter reasoning vision-language model (VLM) for physical AI and robotics developed by Nvidia.

Embodied AINVIDIA

NVLM

NVLM (short for NVIDIA Vision Language Model), released as NVLM 1.0, is a family of open multimodal large language models developed by Nvidia.

Large Language ModelsNVIDIA

PaliGemma

PaliGemma is an open vision-language model developed by Google that pairs the SigLIP image encoder with a Gemma language model, takes an image plus a text prompt as input, and produces text as output.

GoogleOpen Source AI

Paper2Video

Paper2Video (full title: Paper2Video: Automatic Video Generation from Scientific Papers) is a research project from Show Lab at the National University of Singapore that formalizes and evaluates automatic…

AI BenchmarksAI Research

Pika (video generation)

Pika is an artificial intelligence video generation platform developed by Pika Labs, Inc. that lets users create and edit short videos from text prompts, images, and existing clips, and it is best known for…

AI CompaniesAI Models

Pixtral

Pixtral is a family of multimodal vision-language models developed by Mistral AI, a French AI company founded in April 2023.

AI CompaniesAI Models

Project Astra

Project Astra is a research prototype from Google DeepMind that explores what a universal AI assistant might look like: a single agent that can see and hear the world in real time through a device camera and…

AI AgentsGoogle DeepMind

Qwen-VL

Qwen-VL is the first family of open vision-language (multimodal) models from the Qwen team at Alibaba Cloud, able to take images, text, and bounding boxes as input and produce text and bounding boxes as output.

Chinese AIOpen Source AI

Qwen2-VL

Qwen2-VL is a family of open-weight vision-language models released by the Qwen team at Alibaba Cloud between August and September 2024, in 2B, 7B, and 72B Instruct sizes.

Chinese AIOpen Source AI

Qwen2.5-VL

Qwen2.5-VL is a series of open-weight vision-language models released on 26 January 2025 by the Qwen team at Alibaba (Alibaba Cloud), succeeding the earlier Qwen2-VL family.

Chinese AIOpen Source AI

Qwen3-Omni

Qwen3-Omni is a natively end-to-end omni-modal foundation model developed by the Qwen team at Alibaba Cloud, capable of understanding text, images, audio, and video and generating both text and natural speech…

Chinese AILarge Language Models

Reka AI

Reka AI (commonly referred to as Reka) is an artificial intelligence research and product company, founded in 2022 by former Google DeepMind, Google Brain, Meta FAIR, and Baidu researchers, that builds…

AI CompaniesLarge Language Models

Reka Core

Reka Core is a frontier class multimodal foundation model developed by Reka AI, a research and product company founded in 2022 by former scientists from DeepMind, Google Brain, Meta FAIR, and Baidu.

AI ModelsLarge Language Models

Reka Edge

Reka Edge is a 7-billion-parameter multimodal language model developed by Reka AI, introduced in April 2024 as the smallest member of the company's first publicly described model family.

AI ModelsLarge Language Models

Reka Flash

Reka Flash is a family of multimodal large language models developed by Reka AI, a San Francisco Bay Area research company founded in 2022 by former researchers from Google DeepMind, Meta FAIR, and Google.

AI ModelsLarge Language Models

SigLIP

SigLIP (Sigmoid Loss for Language-Image Pre-training) is a family of vision-language encoders developed by researchers at Google DeepMind that pre-trains image and text encoders by treating each image-text…

Computer VisionGoogle DeepMind

Skywork-R1V

Skywork-R1V is an open-weight family of multimodal reasoning models released by Skywork AI, the AGI and AIGC division of Beijing Kunlun Tech Co., Ltd. (Kunlun Wanwei).

Chinese AIReasoning Models

SmolVLA

SmolVLA (Small Vision-Language-Action) is a compact, open-source vision-language-action model (VLA) for robotics developed by Hugging Face and released in June 2025.

AI HardwareAI Models

Sora 2

Sora 2 is a text-to-video and audio generation model developed by OpenAI, released on September 30, 2025, that OpenAI called "the GPT-3.5 moment for video." It succeeded the original Sora research preview from…

AI ModelsGenerative AI

Starchild-1

Starchild-1 is a real-time audio-video world model built by the AI lab Odyssey. It generates synchronized video and sound autoregressively, chunk by chunk, while a user streams new text, speech and action…

AI ModelsVideo Generation

UI-TARS

UI-TARS is a native graphical user interface (GUI) agent model developed by ByteDance through its Seed research team.

AI AgentsChinese AI

Video-MME

Video-MME (Video Multi-Modal Evaluation) is a benchmark for testing how well multimodal large language models (MLLMs) understand video, built from 900 manually selected videos totaling 254 hours and 2,700…

AI BenchmarksComputer Vision

Vision-language-action model

A vision-language-action model (VLA model, or VLA) is a machine-learning model for robotics that uses visual observations and a natural-language instruction to generate actions a robot can execute.

Embodied AIRobotics

WeMM-Embedding

WeMM-Embedding (WeChat Multi-Modal Embedding) is a family of open-weight universal multimodal embedding models built by the WeChat Vision team at Tencent.

AI ModelsChinese AI

ZeroBench

ZeroBench is a visual reasoning benchmark built to be effectively impossible for current frontier large multimodal models, which score 0.0% on its main questions.

AI BenchmarksComputer Vision

data2vec

data2vec is a self-supervised learning framework from Meta AI (then Facebook AI Research) that applies the same training method to three different input types: speech, computer vision, and text.

Machine LearningMeta AI