Multimodal AI

Explore Multimodal AI through related topics and the articles other pages reference most.

Explore articles

Reset filters
Browse subtopics: Open Source AI

Articles that also belong to these categories. Counts cover all of Multimodal AI.

Showing 1-28 of 28 articles

CogVLM

CogVLM is an open vision language model developed by Zhipu AI and the Knowledge Engineering Group (KEG) at Tsinghua University.

Chinese AIOpen Source AI

DeepSeek V4.1-Flash

DeepSeek V4.1-Flash is an open-weight multimodal mixture-of-experts model released by DeepSeek on September 10, 2026. It accepts text and images and generates text.

AI ModelsChinese AI

DeepSeek-OCR

DeepSeek-OCR is an open-source optical character recognition (OCR) and document-understanding system released by DeepSeek on 20 October 2025 that pioneers a contexts optical compression paradigm: it encodes…

Chinese AIComputer Vision

Donut (Model)

Donut (Document understanding transformer) is an OCR-free visual document understanding model introduced by researchers at NAVER CLOVA in the paper "OCR-free Document Understanding Transformer," first posted…

AI ModelsComputer Vision

GLM-5.3-Flash

GLM-5.3-Flash is an open-weight, natively multimodal mixture-of-experts large language model released by Z.ai on August 26, 2026.

AI ModelsChinese AI

Gemma 3

Gemma 3 is a family of open-weight large language models developed by Google DeepMind and released on March 12, 2025.

AI ModelsGoogle

InternVL

InternVL is a family of open-source multimodal large language models developed by the OpenGVLab research group at the Shanghai Artificial Intelligence Laboratory in collaboration with academic partners…

Chinese AIOpen Source AI

Llama 3.2

Llama 3.2 is a family of four open-weight large language models released by Meta on September 25, 2024, comprising lightweight 1 billion and 3 billion parameter text-only models for on-device AI and the 11…

AI ModelsLarge Language Models

MiniCPM-V

MiniCPM-V is a family of open-weights multimodal large language models published by OpenBMB, the shared open-source brand of Tsinghua University's Natural Language Processing lab (THUNLP) and the Beijing…

Chinese AIOpen Source AI

Molmo

Molmo is a family of open-weight, open-data vision-language models (VLMs) released by the Allen Institute for AI (Ai2) on 25 September 2024.

AI ModelsOpen Source AI

Muse Glimmer

Muse Glimmer is an open-weight text-and-image model developed by Meta AI for local agent and coding workloads. Meta released the model on August 10, 2026 under the identifier meta-models/Muse-Glimmer-30B.

AI AgentsAI Models

PaliGemma

PaliGemma is an open vision-language model developed by Google that pairs the SigLIP image encoder with a Gemma language model, takes an image plus a text prompt as input, and produces text as output.

GoogleOpen Source AI

Pixtral

Pixtral is a family of multimodal vision-language models developed by Mistral AI, a French AI company founded in April 2023.

AI CompaniesAI Models

Qwen-VL

Qwen-VL is the first family of open vision-language (multimodal) models from the Qwen team at Alibaba Cloud, able to take images, text, and bounding boxes as input and produce text and bounding boxes as output.

Chinese AIOpen Source AI

Qwen2-VL

Qwen2-VL is a family of open-weight vision-language models released by the Qwen team at Alibaba Cloud between August and September 2024, in 2B, 7B, and 72B Instruct sizes.

Chinese AIOpen Source AI

Qwen2.5-VL

Qwen2.5-VL is a series of open-weight vision-language models released on 26 January 2025 by the Qwen team at Alibaba (Alibaba Cloud), succeeding the earlier Qwen2-VL family.

Chinese AIOpen Source AI

Qwen3-Omni

Qwen3-Omni is a natively end-to-end omni-modal foundation model developed by the Qwen team at Alibaba Cloud, capable of understanding text, images, audio, and video and generating both text and natural speech…

Chinese AILarge Language Models

Reka Flash

Reka Flash is a family of multimodal large language models developed by Reka AI, a San Francisco Bay Area research company founded in 2022 by former researchers from Google DeepMind, Meta FAIR, and Google.

AI ModelsLarge Language Models

SmolVLA

SmolVLA (Small Vision-Language-Action) is a compact, open-source vision-language-action model (VLA) for robotics developed by Hugging Face and released in June 2025.

AI HardwareAI Models

WeMM-Embedding

WeMM-Embedding (WeChat Multi-Modal Embedding) is a family of open-weight universal multimodal embedding models built by the WeChat Vision team at Tencent.

AI ModelsChinese AI