Baidu ERNIE
Baidu ERNIE (Enhanced Representation through Knowledge Integration) is the family of large language and multimodal foundation models built by the Chinese technology company Baidu, spanning the original 2019…
Explore Multimodal AI through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of Multimodal AI.
Showing 1-27 of 27 articles
Baidu ERNIE (Enhanced Representation through Knowledge Integration) is the family of large language and multimodal foundation models built by the Chinese technology company Baidu, spanning the original 2019…
CogAgent is an open visual language model built to act as a graphical user interface (GUI) agent: given a screenshot and a natural-language goal, it predicts the next on-screen action, such as where to click…
CogVLM is an open vision language model developed by Zhipu AI and the Knowledge Engineering Group (KEG) at Tsinghua University.
DeepSeek Janus is a family of open-weight unified multimodal models from Chinese AI lab DeepSeek that perform both image understanding and text-to-image generation in a single autoregressive Transformer.
DeepSeek V4.1-Flash is an open-weight multimodal mixture-of-experts model released by DeepSeek on September 10, 2026. It accepts text and images and generates text.
DeepSeek-OCR is an open-source optical character recognition (OCR) and document-understanding system released by DeepSeek on 20 October 2025 that pioneers a contexts optical compression paradigm: it encodes…
DeepSeek-VL is the first open-source vision-language model series from DeepSeek, the Chinese AI company.
DeepSeek-VL2 is an open-weights family of Mixture-of-Experts (MoE) vision-language models released by the Chinese AI laboratory DeepSeek on December 13, 2024.
Doubao Seed 1.6 is a family of general-purpose foundation models developed by the ByteDance Seed research team and released through Volcano Engine on 11 June 2025 at the company's Force Original Power…
ERNIE 4.5 is a family of large language models released by the Chinese technology company Baidu, open-sourced on June 30, 2025 under the Apache 2.0 license .
ERNIE 5.0 is a natively omni-modal foundation model from Baidu, unveiled at the company's annual Baidu World 2025 conference in Beijing on 13 November 2025 as the flagship in the ERNIE line at launch.
GLM-5.3-Flash is an open-weight, natively multimodal mixture-of-experts large language model released by Z.ai on August 26, 2026.
HY-World 2.0 (also written HunyuanWorld 2.0 or Hunyuan World Model 2.0) is an open multimodal 3D world model from Tencent's Hunyuan team, with a technical report and first code release published on April 16
InternVL is a family of open-source multimodal large language models developed by the OpenGVLab research group at the Shanghai Artificial Intelligence Laboratory in collaboration with academic partners…
MiniCPM-V is a family of open-weights multimodal large language models published by OpenBMB, the shared open-source brand of Tsinghua University's Natural Language Processing lab (THUNLP) and the Beijing…
QvQ (styled QVQ) is a family of experimental visual reasoning models from the Qwen team at Alibaba.
Qwen-VL is the first family of open vision-language (multimodal) models from the Qwen team at Alibaba Cloud, able to take images, text, and bounding boxes as input and produce text and bounding boxes as output.
Qwen2-Audio is an audio-language model developed by the Qwen team at Alibaba Cloud, released in August 2024 .
Qwen2-VL is a family of open-weight vision-language models released by the Qwen team at Alibaba Cloud between August and September 2024, in 2B, 7B, and 72B Instruct sizes.
Qwen2.5-VL is a series of open-weight vision-language models released on 26 January 2025 by the Qwen team at Alibaba (Alibaba Cloud), succeeding the earlier Qwen2-VL family.
Qwen3-Omni is a natively end-to-end omni-modal foundation model developed by the Qwen team at Alibaba Cloud, capable of understanding text, images, audio, and video and generating both text and natural speech…
Qwen3-VL is a family of open-weight vision-language models built by the Qwen team at Alibaba Cloud, first released in September 2025.
Qwen3.8-Flash-Next is an experimental open-weight multimodal large language model released by Alibaba Group's Qwen team on August 26, 2026.
Seedream 4.0 is a unified image generation and editing model built by the Seed team at ByteDance.
Skywork-R1V is an open-weight family of multimodal reasoning models released by Skywork AI, the AGI and AIGC division of Beijing Kunlun Tech Co., Ltd. (Kunlun Wanwei).
UI-TARS is a native graphical user interface (GUI) agent model developed by ByteDance through its Seed research team.
WeMM-Embedding (WeChat Multi-Modal Embedding) is a family of open-weight universal multimodal embedding models built by the WeChat Vision team at Tencent.