Llama 3.2
Llama 3.2 is a family of four open-weight large language models released by Meta on September 25, 2024, comprising lightweight 1 billion and 3 billion parameter text-only models for on-device AI and the 11…
Explore Multimodal AI through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of Multimodal AI.
Showing 61-118 of 118 articles
Llama 3.2 is a family of four open-weight large language models released by Meta on September 25, 2024, comprising lightweight 1 billion and 3 billion parameter text-only models for on-device AI and the 11…
Llama 3.2 Vision is the set of multimodal (image-plus-text) models in Meta's Llama 3.2 family, released on September 25, 2024 at the Meta Connect 2024 developer conference.
Llama 4 Scout and Llama 4 Maverick are open-weight, natively multimodal AI large language models developed by Meta and released on April 5, 2025.
MAGI-2 Preview is a public research release of a unified audio-video generation model developed by Sand.ai.
MM-BrowseComp is a benchmark for evaluating multimodal web-browsing AI agents, introduced in August 2025 by researchers from ByteDance, Nanjing University, M-A-P, the Institute of Automation of the Chinese…
MMMU (Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark) is a multimodal AI benchmark of 11,550 college-level questions that pairs text with images to test expert knowledge and…
MMMU-Pro is a rigorous benchmark for evaluating multimodal AI systems on college-level, expert questions that genuinely require seeing an image, built as a harder and more robust version of the original MMMU…
MMStar (Multi-modal Star) is a vision-language model evaluation benchmark consisting of 1,500 multimodal samples that were filtered from six pre-existing benchmarks to ensure both visual dependency (questions…
MathVista is a benchmark for evaluating the mathematical reasoning capabilities of foundation models in visual contexts.
MetaCLIP (Metadata-Curated Language-Image Pre-training) is a data curation recipe and a family of vision-language models from Meta AI, introduced in the 2023 paper "Demystifying CLIP Data" by Hu Xu, Saining…
MiniCPM-V is a family of open-weights multimodal large language models published by OpenBMB, the shared open-source brand of Tsinghua University's Natural Language Processing lab (THUNLP) and the Beijing…
Molmo is a family of open-weight, open-data vision-language models (VLMs) released by the Allen Institute for AI (Ai2) on 25 September 2024.
Muse Glimmer is an open-weight text-and-image model developed by Meta AI for local agent and coding workloads. Meta released the model on August 10, 2026 under the identifier meta-models/Muse-Glimmer-30B.
Muse Image is a proprietary image-generation and editing system developed by Meta Superintelligence Labs, a division of Meta.
Muse Spark is a proprietary multimodal reasoning model developed by Meta Superintelligence Labs (MSL), the artificial intelligence division Meta reorganized in 2025.
The original NVIDIA Cosmos Reason release is an open, customizable, 7-billion-parameter reasoning vision-language model (VLM) for physical AI and robotics developed by Nvidia.
NVLM (short for NVIDIA Vision Language Model), released as NVLM 1.0, is a family of open multimodal large language models developed by Nvidia.
Nano Banana Pro is a professional grade image generation and editing model from Google DeepMind, released on November 20, 2025.
PaLM-E (short for Pathways Language Model, Embodied) is an embodied multimodal large language model introduced by Google and TU Berlin in March 2023 that injects continuous sensor and image observations…
PaliGemma is an open vision-language model developed by Google that pairs the SigLIP image encoder with a Gemma language model, takes an image plus a text prompt as input, and produces text as output.
Paper2Video (full title: Paper2Video: Automatic Video Generation from Scientific Papers) is a research project from Show Lab at the National University of Singapore that formalizes and evaluates automatic…
Perception Encoder (PE) is a family of vision and vision-language encoders from Meta AI's Fundamental AI Research (FAIR) group, released in April 2025
Pika is an artificial intelligence video generation platform developed by Pika Labs, Inc. that lets users create and edit short videos from text prompts, images, and existing clips, and it is best known for…
Pixtral is a family of multimodal vision-language models developed by Mistral AI, a French AI company founded in April 2023.
Project Astra is a research prototype from Google DeepMind that explores what a universal AI assistant might look like: a single agent that can see and hear the world in real time through a device camera and…
QvQ (styled QVQ) is a family of experimental visual reasoning models from the Qwen team at Alibaba.
Qwen-VL is the first family of open vision-language (multimodal) models from the Qwen team at Alibaba Cloud, able to take images, text, and bounding boxes as input and produce text and bounding boxes as output.
Qwen2-Audio is an audio-language model developed by the Qwen team at Alibaba Cloud, released in August 2024 .
Qwen2-VL is a family of open-weight vision-language models released by the Qwen team at Alibaba Cloud between August and September 2024, in 2B, 7B, and 72B Instruct sizes.
Qwen2.5-VL is a series of open-weight vision-language models released on 26 January 2025 by the Qwen team at Alibaba (Alibaba Cloud), succeeding the earlier Qwen2-VL family.
Qwen3-Omni is a natively end-to-end omni-modal foundation model developed by the Qwen team at Alibaba Cloud, capable of understanding text, images, audio, and video and generating both text and natural speech…
Qwen3-VL is a family of open-weight vision-language models built by the Qwen team at Alibaba Cloud, first released in September 2025.
Qwen3.8-Flash-Next is an experimental open-weight multimodal large language model released by Alibaba Group's Qwen team on August 26, 2026.
Reka AI (commonly referred to as Reka) is an artificial intelligence research and product company, founded in 2022 by former Google DeepMind, Google Brain, Meta FAIR, and Baidu researchers, that builds…
Reka Core is a frontier class multimodal foundation model developed by Reka AI, a research and product company founded in 2022 by former scientists from DeepMind, Google Brain, Meta FAIR, and Baidu.
Reka Edge is a 7-billion-parameter multimodal language model developed by Reka AI, introduced in April 2024 as the smallest member of the company's first publicly described model family.
Reka Flash is a family of multimodal large language models developed by Reka AI, a San Francisco Bay Area research company founded in 2022 by former researchers from Google DeepMind, Meta FAIR, and Google.
Robostral Navigate is an 8-billion-parameter vision language model developed by Mistral AI for instruction-following robot navigation.
SWE-bench Multimodal (also written SWE-bench M) is a benchmark that measures whether autonomous software-engineering systems can resolve bugs in visual
Seedream 4.0 is a unified image generation and editing model built by the Seed team at ByteDance.
SigLIP (Sigmoid Loss for Language-Image Pre-training) is a family of vision-language encoders developed by researchers at Google DeepMind that pre-trains image and text encoders by treating each image-text…
Skywork-R1V is an open-weight family of multimodal reasoning models released by Skywork AI, the AGI and AIGC division of Beijing Kunlun Tech Co., Ltd. (Kunlun Wanwei).
SmolVLA (Small Vision-Language-Action) is a compact, open-source vision-language-action model (VLA) for robotics developed by Hugging Face and released in June 2025.
Sora 2 is a text-to-video and audio generation model developed by OpenAI, released on September 30, 2025, that OpenAI called "the GPT-3.5 moment for video." It succeeded the original Sora research preview from…
Starchild-1 is a real-time audio-video world model built by the AI lab Odyssey. It generates synchronized video and sound autoregressively, chunk by chunk, while a user streams new text, speech and action…
T-Rex: Tactile-Reactive Dexterous Manipulation is a 2026 robot learning research system for contact-rich, bimanual manipulation.
Text-to-image models are generative artificial intelligence systems that synthesize a new image from a natural-language description, called a prompt.
UI-TARS is a native graphical user interface (GUI) agent model developed by ByteDance through its Seed research team.
Video-MME (Video Multi-Modal Evaluation) is a benchmark for testing how well multimodal large language models (MLLMs) understand video, built from 900 manually selected videos totaling 254 hours and 2,700…
A vision encoder is the neural network component that turns an image into a sequence of numerical vectors that other models can consume.
Vision-language models (VLMs) are artificial intelligence models that learn relationships between visual data and natural language.
A vision-language-action model (VLA model, or VLA) is a machine-learning model for robotics that uses visual observations and a natural-language instruction to generate actions a robot can execute.
Visual question answering models are AI systems that take an image and a natural language question about that image and return a natural language answer.
Visual question answering (VQA) is the task of producing a natural-language answer to a natural-language question about an image.
WeMM-Embedding (WeChat Multi-Modal Embedding) is a family of open-weight universal multimodal embedding models built by the WeChat Vision team at Tencent.
ZeroBench is a visual reasoning benchmark built to be effectively impossible for current frontier large multimodal models, which score 0.0% on its main questions.
data2vec is a self-supervised learning framework from Meta AI (then Facebook AI Research) that applies the same training method to three different input types: speech, computer vision, and text.
dots3-note Preview is an open-weight multimodal model and large language model developed by dots studio, an AI team at Xiaohongshu.