CogVLM
CogVLM is an open vision language model developed by Zhipu AI and the Knowledge Engineering Group (KEG) at Tsinghua University.
Explore Multimodal AI through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of Multimodal AI.
Showing 1-28 of 28 articles
CogVLM is an open vision language model developed by Zhipu AI and the Knowledge Engineering Group (KEG) at Tsinghua University.
DeepSeek V4.1-Flash is an open-weight multimodal mixture-of-experts model released by DeepSeek on September 10, 2026. It accepts text and images and generates text.
DeepSeek-OCR is an open-source optical character recognition (OCR) and document-understanding system released by DeepSeek on 20 October 2025 that pioneers a contexts optical compression paradigm: it encodes…
DiffusionGemma is an experimental open-weight large language model developed by Google DeepMind for multimodal text generation through discrete diffusion.
Donut (Document understanding transformer) is an OCR-free visual document understanding model introduced by researchers at NAVER CLOVA in the paper "OCR-free Document Understanding Transformer," first posted…
ERNIE 4.5 is a family of large language models released by the Chinese technology company Baidu, open-sourced on June 30, 2025 under the Apache 2.0 license .
GLM-5.3-Flash is an open-weight, natively multimodal mixture-of-experts large language model released by Z.ai on August 26, 2026.
Gemma 3 is a family of open-weight large language models developed by Google DeepMind and released on March 12, 2025.
InternVL is a family of open-source multimodal large language models developed by the OpenGVLab research group at the Shanghai Artificial Intelligence Laboratory in collaboration with academic partners…
LLaVA (Large Language and Vision Assistant) is an open-source family of multimodal large language models that connects a frozen vision encoder to a pre-trained large language model through a single learned…
Llama 3.2 is a family of four open-weight large language models released by Meta on September 25, 2024, comprising lightweight 1 billion and 3 billion parameter text-only models for on-device AI and the 11…
Llama 4 Scout and Llama 4 Maverick are open-weight, natively multimodal AI large language models developed by Meta and released on April 5, 2025.
MAGI-2 Preview is a public research release of a unified audio-video generation model developed by Sand.ai.
MiniCPM-V is a family of open-weights multimodal large language models published by OpenBMB, the shared open-source brand of Tsinghua University's Natural Language Processing lab (THUNLP) and the Beijing…
Molmo is a family of open-weight, open-data vision-language models (VLMs) released by the Allen Institute for AI (Ai2) on 25 September 2024.
Muse Glimmer is an open-weight text-and-image model developed by Meta AI for local agent and coding workloads. Meta released the model on August 10, 2026 under the identifier meta-models/Muse-Glimmer-30B.
PaliGemma is an open vision-language model developed by Google that pairs the SigLIP image encoder with a Gemma language model, takes an image plus a text prompt as input, and produces text as output.
Pixtral is a family of multimodal vision-language models developed by Mistral AI, a French AI company founded in April 2023.
Qwen-VL is the first family of open vision-language (multimodal) models from the Qwen team at Alibaba Cloud, able to take images, text, and bounding boxes as input and produce text and bounding boxes as output.
Qwen2-VL is a family of open-weight vision-language models released by the Qwen team at Alibaba Cloud between August and September 2024, in 2B, 7B, and 72B Instruct sizes.
Qwen2.5-VL is a series of open-weight vision-language models released on 26 January 2025 by the Qwen team at Alibaba (Alibaba Cloud), succeeding the earlier Qwen2-VL family.
Qwen3-Omni is a natively end-to-end omni-modal foundation model developed by the Qwen team at Alibaba Cloud, capable of understanding text, images, audio, and video and generating both text and natural speech…
Qwen3-VL is a family of open-weight vision-language models built by the Qwen team at Alibaba Cloud, first released in September 2025.
Qwen3.8-Flash-Next is an experimental open-weight multimodal large language model released by Alibaba Group's Qwen team on August 26, 2026.
Reka Flash is a family of multimodal large language models developed by Reka AI, a San Francisco Bay Area research company founded in 2022 by former researchers from Google DeepMind, Meta FAIR, and Google.
SmolVLA (Small Vision-Language-Action) is a compact, open-source vision-language-action model (VLA) for robotics developed by Hugging Face and released in June 2025.
WeMM-Embedding (WeChat Multi-Modal Embedding) is a family of open-weight universal multimodal embedding models built by the WeChat Vision team at Tencent.
dots3-note Preview is an open-weight multimodal model and large language model developed by dots studio, an AI team at Xiaohongshu.