Baidu ERNIE
Baidu ERNIE (Enhanced Representation through Knowledge Integration) is the family of large language and multimodal foundation models built by the Chinese technology company Baidu, spanning the original 2019…
Explore Multimodal AI through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of Multimodal AI.
Showing 1-41 of 41 articles
Baidu ERNIE (Enhanced Representation through Knowledge Integration) is the family of large language and multimodal foundation models built by the Chinese technology company Baidu, spanning the original 2019…
Claude Sonnet 4.5 is a multimodal large language model (LLM) developed by Anthropic and released on September 29, 2025, which Anthropic described at launch as "the best coding model in the world." It is a…
DeepSeek V4.1-Flash is an open-weight multimodal mixture-of-experts model released by DeepSeek on September 10, 2026. It accepts text and images and generates text.
DiffusionGemma is an experimental open-weight large language model developed by Google DeepMind for multimodal text generation through discrete diffusion.
Doubao Seed 1.6 is a family of general-purpose foundation models developed by the ByteDance Seed research team and released through Volcano Engine on 11 June 2025 at the company's Force Original Power…
ERNIE 4.5 is a family of large language models released by the Chinese technology company Baidu, open-sourced on June 30, 2025 under the Apache 2.0 license .
ERNIE 5.0 is a natively omni-modal foundation model from Baidu, unveiled at the company's annual Baidu World 2025 conference in Beijing on 13 November 2025 as the flagship in the ERNIE line at launch.
GLM-5.3-Flash is an open-weight, natively multimodal mixture-of-experts large language model released by Z.ai on August 26, 2026.
GPT-4V, also written GPT-4V(ision) and read as "GPT-4 with vision," is the image-understanding capability that OpenAI added to its GPT-4 large language model, letting a user supply one or more images alongside…
Gemini 1.0 is the first generation of Gemini, the family of natively multimodal AI models that Google DeepMind announced on 6 December 2023 .
Gemini 1.5 Flash is a lightweight, low-latency multimodal large language model from Google DeepMind, released at Google I/O on May 14, 2024, as the fast and cost-efficient member of the Gemini 1.5 family.
Gemini 1.5 Pro is a multimodal large language model developed by Google DeepMind and announced on February 15, 2024, as the flagship model of the Gemini 1.5 generation.
Gemini 2.0 Flash is a fast, low-cost multimodal large language model built by Google DeepMind as the flagship workhorse of the Gemini 2.0 generation, designed for the agentic era with native tool use, a 1…
Gemini 2.0 Flash-Lite is a large language model developed by Google DeepMind and released as the most cost-efficient member of the Gemini 2.0 model family.
Gemini 2.5 Flash is a fast, cost-optimized multimodal large language model developed by Google DeepMind and the mid-tier member of the Gemini 2.5 family.
Gemini 3 is the third major generation of the Gemini family of multimodal models from Google DeepMind, launched on November 18, 2025 with Gemini 3 Pro as the flagship and described by Google as "our most…
Gemini 3 Flash is a multimodal large language model released by Google on December 17, 2025 as the fast, lower-cost sibling to Gemini 3 Pro in the Gemini 3 family.
Gemini 3.1 Pro is a large language model developed by Google DeepMind and released on 19 February 2026 as a point-release upgrade to Gemini 3 Pro .
Gemini 3.5 Flash is a fast frontier large language model developed by Google DeepMind, announced at Google I/O 2026 on May 19, 2026 and made generally available the same day .
Gemini 3.6 Flash is a proprietary, multimodal large language model released by Google on July 21, 2026. It belongs to the Gemini 3 series and uses the stable API identifier gemini-3.6-flash.
Gemini 3.7 Flash is a proprietary, multimodal large language model released by Google on August 13, 2026.
Gemini 3.8 Flash is a multimodal model in Google's Gemini family. Google DeepMind released it on September 2, 2026 as a generally available model for software engineering, tool-using agents, and knowledge work.
Gemini Ultra (branded Ultra 1.0) was the largest and most capable model in the Gemini 1.0 family, the first generation of natively multimodal large language models from Google DeepMind.
Gemma 3 is a family of open-weight large language models developed by Google DeepMind and released on March 12, 2025.
LayoutLM is a family of pre-trained multimodal models developed by Microsoft Research for document AI, the task of automatically reading and understanding visually rich documents such as forms, invoices…
Llama 3.2 is a family of four open-weight large language models released by Meta on September 25, 2024, comprising lightweight 1 billion and 3 billion parameter text-only models for on-device AI and the 11…
Llama 3.2 Vision is the set of multimodal (image-plus-text) models in Meta's Llama 3.2 family, released on September 25, 2024 at the Meta Connect 2024 developer conference.
Llama 4 Scout and Llama 4 Maverick are open-weight, natively multimodal AI large language models developed by Meta and released on April 5, 2025.
MMMU-Pro is a rigorous benchmark for evaluating multimodal AI systems on college-level, expert questions that genuinely require seeing an image, built as a harder and more robust version of the original MMMU…
Muse Glimmer is an open-weight text-and-image model developed by Meta AI for local agent and coding workloads. Meta released the model on August 10, 2026 under the identifier meta-models/Muse-Glimmer-30B.
NVLM (short for NVIDIA Vision Language Model), released as NVLM 1.0, is a family of open multimodal large language models developed by Nvidia.
Pixtral is a family of multimodal vision-language models developed by Mistral AI, a French AI company founded in April 2023.
Qwen3-Omni is a natively end-to-end omni-modal foundation model developed by the Qwen team at Alibaba Cloud, capable of understanding text, images, audio, and video and generating both text and natural speech…
Qwen3-VL is a family of open-weight vision-language models built by the Qwen team at Alibaba Cloud, first released in September 2025.
Qwen3.8-Flash-Next is an experimental open-weight multimodal large language model released by Alibaba Group's Qwen team on August 26, 2026.
Reka AI (commonly referred to as Reka) is an artificial intelligence research and product company, founded in 2022 by former Google DeepMind, Google Brain, Meta FAIR, and Baidu researchers, that builds…
Reka Core is a frontier class multimodal foundation model developed by Reka AI, a research and product company founded in 2022 by former scientists from DeepMind, Google Brain, Meta FAIR, and Baidu.
Reka Edge is a 7-billion-parameter multimodal language model developed by Reka AI, introduced in April 2024 as the smallest member of the company's first publicly described model family.
Reka Flash is a family of multimodal large language models developed by Reka AI, a San Francisco Bay Area research company founded in 2022 by former researchers from Google DeepMind, Meta FAIR, and Google.
Vision-language models (VLMs) are artificial intelligence models that learn relationships between visual data and natural language.
dots3-note Preview is an open-weight multimodal model and large language model developed by dots studio, an AI team at Xiaohongshu.