Multimodal AI

Explore Multimodal AI through related topics and the articles other pages reference most.

Explore articles

Reset filters
Browse subtopics: AI Models

Articles that also belong to these categories. Counts cover all of Multimodal AI.

Showing 1-45 of 45 articles

Chameleon (Meta AI)

Chameleon is a family of early-fusion, token-based mixed-modal foundation models from Meta AI's Fundamental AI Research (FAIR) group that represents both images and text as discrete tokens in a single unified…

AI ModelsMeta AI

DeepSeek Janus

DeepSeek Janus is a family of open-weight unified multimodal models from Chinese AI lab DeepSeek that perform both image understanding and text-to-image generation in a single autoregressive Transformer.

AI ModelsChinese AI

DeepSeek V4.1-Flash

DeepSeek V4.1-Flash is an open-weight multimodal mixture-of-experts model released by DeepSeek on September 10, 2026. It accepts text and images and generates text.

AI ModelsChinese AI

Document Question Answering Models

Document question answering models (DocQA, sometimes called DocVQA for document visual question answering) are machine learning systems that take a document image or PDF together with a natural language…

AI Models

Donut (Model)

Donut (Document understanding transformer) is an OCR-free visual document understanding model introduced by researchers at NAVER CLOVA in the paper "OCR-free Document Understanding Transformer," first posted…

AI ModelsComputer Vision

Doubao Seed 1.6

Doubao Seed 1.6 is a family of general-purpose foundation models developed by the ByteDance Seed research team and released through Volcano Engine on 11 June 2025 at the company's Force Original Power…

AI ModelsChinese AI

Feature Extraction Models

Feature extraction models are machine learning systems that transform raw inputs such as text, images, or audio into dense numerical vectors known as embeddings or hidden-state representations.

AI Models

GLM-5.3-Flash

GLM-5.3-Flash is an open-weight, natively multimodal mixture-of-experts large language model released by Z.ai on August 26, 2026.

AI ModelsChinese AI

GPT Image 1

GPT Image 1 (API identifier gpt-image-1) is a natively multimodal image generation model developed by OpenAI, integrated into ChatGPT on March 25, 2025, and released as a standalone API on April 23, 2025.

AI ModelsGenerative AI

Gemini 3

Gemini 3 is the third major generation of the Gemini family of multimodal models from Google DeepMind, launched on November 18, 2025 with Gemini 3 Pro as the flagship and described by Google as "our most…

AI ModelsGoogle

Gemini 3 Flash

Gemini 3 Flash is a multimodal large language model released by Google on December 17, 2025 as the fast, lower-cost sibling to Gemini 3 Pro in the Gemini 3 family.

AI ModelsGoogle

Gemini 3 Pro

Gemini 3 Pro is the flagship preview model in Google DeepMind's Gemini 3 family of multimodal models, launched on November 18, 2025 as what Google called "our most intelligent model." It became the first…

AI ModelsGoogle

Gemini 3.6 Flash

Gemini 3.6 Flash is a proprietary, multimodal large language model released by Google on July 21, 2026. It belongs to the Gemini 3 series and uses the stable API identifier gemini-3.6-flash.

AI ModelsGenerative AI

Gemini 3.8 Flash

Gemini 3.8 Flash is a multimodal model in Google's Gemini family. Google DeepMind released it on September 2, 2026 as a generally available model for software engineering, tool-using agents, and knowledge work.

AI ModelsGenerative AI

Gemini Robotics 2

Gemini Robotics 2 is a family of three robotics models announced by Google DeepMind on July 30, 2026: a vision-language-action (VLA) model of the same name, an embodied reasoning model called Gemini Robotics…

AI ModelsEmbodied AI

Gemma 3

Gemma 3 is a family of open-weight large language models developed by Google DeepMind and released on March 12, 2025.

AI ModelsGoogle

ImageBind

ImageBind is a multimodal model from Meta AI (its Fundamental AI Research lab) that learns a single joint embedding space across six different modalities: images and video, text, audio, depth, thermal…

AI ModelsMeta AI

Llama 3.2

Llama 3.2 is a family of four open-weight large language models released by Meta on September 25, 2024, comprising lightweight 1 billion and 3 billion parameter text-only models for on-device AI and the 11…

AI ModelsLarge Language Models

Molmo

Molmo is a family of open-weight, open-data vision-language models (VLMs) released by the Allen Institute for AI (Ai2) on 25 September 2024.

AI ModelsOpen Source AI

Muse Glimmer

Muse Glimmer is an open-weight text-and-image model developed by Meta AI for local agent and coding workloads. Meta released the model on August 10, 2026 under the identifier meta-models/Muse-Glimmer-30B.

AI AgentsAI Models

Pika (video generation)

Pika is an artificial intelligence video generation platform developed by Pika Labs, Inc. that lets users create and edit short videos from text prompts, images, and existing clips, and it is best known for…

AI CompaniesAI Models

Pixtral

Pixtral is a family of multimodal vision-language models developed by Mistral AI, a French AI company founded in April 2023.

AI CompaniesAI Models

Reka Core

Reka Core is a frontier class multimodal foundation model developed by Reka AI, a research and product company founded in 2022 by former scientists from DeepMind, Google Brain, Meta FAIR, and Baidu.

AI ModelsLarge Language Models

Reka Edge

Reka Edge is a 7-billion-parameter multimodal language model developed by Reka AI, introduced in April 2024 as the smallest member of the company's first publicly described model family.

AI ModelsLarge Language Models

Reka Flash

Reka Flash is a family of multimodal large language models developed by Reka AI, a San Francisco Bay Area research company founded in 2022 by former researchers from Google DeepMind, Meta FAIR, and Google.

AI ModelsLarge Language Models

SmolVLA

SmolVLA (Small Vision-Language-Action) is a compact, open-source vision-language-action model (VLA) for robotics developed by Hugging Face and released in June 2025.

AI HardwareAI Models

Sora 2

Sora 2 is a text-to-video and audio generation model developed by OpenAI, released on September 30, 2025, that OpenAI called "the GPT-3.5 moment for video." It succeeded the original Sora research preview from…

AI ModelsGenerative AI

Starchild-1

Starchild-1 is a real-time audio-video world model built by the AI lab Odyssey. It generates synchronized video and sound autoregressively, chunk by chunk, while a user streams new text, speech and action…

AI ModelsVideo Generation

WeMM-Embedding

WeMM-Embedding (WeChat Multi-Modal Embedding) is a family of open-weight universal multimodal embedding models built by the WeChat Vision team at Tencent.

AI ModelsChinese AI