Multimodal AI

Explore Multimodal AI through related topics and the articles other pages reference most.

Most referenced in this topic

Ranked by links from other AI Wiki pages.

Explore articles

Browse subtopics (42)

Articles that also belong to these categories. Counts cover all of Multimodal AI.

Showing 1-60 of 118 articles

Baidu ERNIE

Baidu ERNIE (Enhanced Representation through Knowledge Integration) is the family of large language and multimodal foundation models built by the Chinese technology company Baidu, spanning the original 2019…

Chinese AILarge Language Models

CLIP Score

CLIP Score (also written CLIPScore or CLIP-S) is a reference-free automatic evaluation metric that measures how well a text caption matches an image, computed as the rescaled cosine similarity of the image and…

AI BenchmarksComputer Vision

CM3leon

CM3leon (pronounced "chameleon") is a multimodal generative model from Meta AI, introduced in July 2023, that handles both text-to-image and image-to-text generation in a single architecture.

Image GenerationMeta AI

Chameleon (Meta AI)

Chameleon is a family of early-fusion, token-based mixed-modal foundation models from Meta AI's Fundamental AI Research (FAIR) group that represents both images and text as discrete tokens in a single unified…

AI ModelsMeta AI

CogAgent

CogAgent is an open visual language model built to act as a graphical user interface (GUI) agent: given a screenshot and a natural-language goal, it predicts the next on-screen action, such as where to click…

AI AgentsChinese AI

CogVLM

CogVLM is an open vision language model developed by Zhipu AI and the Knowledge Engineering Group (KEG) at Tsinghua University.

Chinese AIOpen Source AI

Computer-use agent

A computer-use agent (CUA) is a category of AI agent in artificial intelligence that performs tasks by directly operating a general-purpose computer's graphical user interface (GUI) the way a human does, by…

AI AgentsArtificial Intelligence

DeepSeek Janus

DeepSeek Janus is a family of open-weight unified multimodal models from Chinese AI lab DeepSeek that perform both image understanding and text-to-image generation in a single autoregressive Transformer.

AI ModelsChinese AI

DeepSeek V4.1-Flash

DeepSeek V4.1-Flash is an open-weight multimodal mixture-of-experts model released by DeepSeek on September 10, 2026. It accepts text and images and generates text.

AI ModelsChinese AI

DeepSeek-OCR

DeepSeek-OCR is an open-source optical character recognition (OCR) and document-understanding system released by DeepSeek on 20 October 2025 that pioneers a contexts optical compression paradigm: it encodes…

Chinese AIComputer Vision

Document Question Answering Models

Document question answering models (DocQA, sometimes called DocVQA for document visual question answering) are machine learning systems that take a document image or PDF together with a natural language…

AI Models

Donut (Model)

Donut (Document understanding transformer) is an OCR-free visual document understanding model introduced by researchers at NAVER CLOVA in the paper "OCR-free Document Understanding Transformer," first posted…

AI ModelsComputer Vision

Doubao Seed 1.6

Doubao Seed 1.6 is a family of general-purpose foundation models developed by the ByteDance Seed research team and released through Volcano Engine on 11 June 2025 at the company's Force Original Power…

AI ModelsChinese AI

ERNIE 5.0

ERNIE 5.0 is a natively omni-modal foundation model from Baidu, unveiled at the company's annual Baidu World 2025 conference in Beijing on 13 November 2025 as the flagship in the ERNIE line at launch.

Chinese AILarge Language Models

ERQA

ERQA (Embodied Reasoning Question Answering) is a multimodal benchmark released by Google DeepMind in March 2025 to evaluate the embodied reasoning capabilities of vision-language models (VLMs) on robotics…

AI BenchmarksEmbodied AI

EgoSchema

EgoSchema is a diagnostic benchmark for evaluating very long-form video language understanding, introduced by Karttikeya Mangalam, Raiymbek Akshulakov, and Jitendra Malik at UC Berkeley.

AI BenchmarksComputer Vision

Feature Extraction Models

Feature extraction models are machine learning systems that transform raw inputs such as text, images, or audio into dense numerical vectors known as embeddings or hidden-state representations.

AI Models

GLM-5.3-Flash

GLM-5.3-Flash is an open-weight, natively multimodal mixture-of-experts large language model released by Z.ai on August 26, 2026.

AI ModelsChinese AI

GPT Image 1

GPT Image 1 (API identifier gpt-image-1) is a natively multimodal image generation model developed by OpenAI, integrated into ChatGPT on March 25, 2025, and released as a standalone API on April 23, 2025.

AI ModelsGenerative AI

GPT-4V (Vision)

GPT-4V, also written GPT-4V(ision) and read as "GPT-4 with vision," is the image-understanding capability that OpenAI added to its GPT-4 large language model, letting a user supply one or more images alongside…

Large Language ModelsOpenAI

GPT-4o mini

GPT-4o mini is a small, low-cost multimodal large language model developed by OpenAI and released on July 18, 2024, as the company's most cost-efficient model at the time.

OpenAISmall Language Models

Gemini 3

Gemini 3 is the third major generation of the Gemini family of multimodal models from Google DeepMind, launched on November 18, 2025 with Gemini 3 Pro as the flagship and described by Google as "our most…

AI ModelsGoogle

Gemini 3 Flash

Gemini 3 Flash is a multimodal large language model released by Google on December 17, 2025 as the fast, lower-cost sibling to Gemini 3 Pro in the Gemini 3 family.

AI ModelsGoogle

Gemini 3 Pro

Gemini 3 Pro is the flagship preview model in Google DeepMind's Gemini 3 family of multimodal models, launched on November 18, 2025 as what Google called "our most intelligent model." It became the first…

AI ModelsGoogle

Gemini 3.6 Flash

Gemini 3.6 Flash is a proprietary, multimodal large language model released by Google on July 21, 2026. It belongs to the Gemini 3 series and uses the stable API identifier gemini-3.6-flash.

AI ModelsGenerative AI

Gemini 3.8 Flash

Gemini 3.8 Flash is a multimodal model in Google's Gemini family. Google DeepMind released it on September 2, 2026 as a generally available model for software engineering, tool-using agents, and knowledge work.

AI ModelsGenerative AI

Gemini Robotics 2

Gemini Robotics 2 is a family of three robotics models announced by Google DeepMind on July 30, 2026: a vision-language-action (VLA) model of the same name, an embodied reasoning model called Gemini Robotics…

AI ModelsEmbodied AI

Gemma 3

Gemma 3 is a family of open-weight large language models developed by Google DeepMind and released on March 12, 2025.

AI ModelsGoogle

HY-World 2.0

HY-World 2.0 (also written HunyuanWorld 2.0 or Hunyuan World Model 2.0) is an open multimodal 3D world model from Tencent's Hunyuan team, with a technical report and first code release published on April 16

Chinese AIWorld Models

HuggingGPT

HuggingGPT is an agent system that uses a large language model as a controller to plan tasks and orchestrate specialist machine learning models hosted on Hugging Face.

AI AgentsAI Research

ImageBind

ImageBind is a multimodal model from Meta AI (its Fundamental AI Research lab) that learns a single joint embedding space across six different modalities: images and video, text, audio, depth, thermal…

AI ModelsMeta AI

InternVL

InternVL is a family of open-source multimodal large language models developed by the OpenGVLab research group at the Shanghai Artificial Intelligence Laboratory in collaboration with academic partners…

Chinese AIOpen Source AI

LayoutLM

LayoutLM is a family of pre-trained multimodal models developed by Microsoft Research for document AI, the task of automatically reading and understanding visually rich documents such as forms, invoices…

Large Language Models