Chameleon (Meta AI)
Chameleon is a family of early-fusion, token-based mixed-modal foundation models from Meta AI's Fundamental AI Research (FAIR) group that represents both images and text as discrete tokens in a single unified…
Explore AI Models through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of AI Models.
Showing 1-44 of 44 articles
Chameleon is a family of early-fusion, token-based mixed-modal foundation models from Meta AI's Fundamental AI Research (FAIR) group that represents both images and text as discrete tokens in a single unified…
DeepSeek Janus is a family of open-weight unified multimodal models from Chinese AI lab DeepSeek that perform both image understanding and text-to-image generation in a single autoregressive Transformer.
DeepSeek V4.1-Flash is an open-weight multimodal mixture-of-experts model released by DeepSeek on September 10, 2026. It accepts text and images and generates text.
DeepSeek-VL is the first open-source vision-language model series from DeepSeek, the Chinese AI company.
Document question answering models (DocQA, sometimes called DocVQA for document visual question answering) are machine learning systems that take a document image or PDF together with a natural language…
Donut (Document understanding transformer) is an OCR-free visual document understanding model introduced by researchers at NAVER CLOVA in the paper "OCR-free Document Understanding Transformer," first posted…
Doubao Seed 1.6 is a family of general-purpose foundation models developed by the ByteDance Seed research team and released through Volcano Engine on 11 June 2025 at the company's Force Original Power…
Feature extraction models are machine learning systems that transform raw inputs such as text, images, or audio into dense numerical vectors known as embeddings or hidden-state representations.
Flamingo is a family of visual language models (VLMs) built by DeepMind and introduced in April 2022 that brought few-shot, in-context learning to multimodal inputs.
GLM-5.3-Flash is an open-weight, natively multimodal mixture-of-experts large language model released by Z.ai on August 26, 2026.
GPT Image 1 (API identifier gpt-image-1) is a natively multimodal image generation model developed by OpenAI, integrated into ChatGPT on March 25, 2025, and released as a standalone API on April 23, 2025.
Gemini 2.5 Flash is a fast, cost-optimized multimodal large language model developed by Google DeepMind and the mid-tier member of the Gemini 2.5 family.
Gemini 3 is the third major generation of the Gemini family of multimodal models from Google DeepMind, launched on November 18, 2025 with Gemini 3 Pro as the flagship and described by Google as "our most…
Gemini 3 Flash is a multimodal large language model released by Google on December 17, 2025 as the fast, lower-cost sibling to Gemini 3 Pro in the Gemini 3 family.
Gemini 3 Pro is the flagship preview model in Google DeepMind's Gemini 3 family of multimodal models, launched on November 18, 2025 as what Google called "our most intelligent model." It became the first…
Gemini 3.6 Flash is a proprietary, multimodal large language model released by Google on July 21, 2026. It belongs to the Gemini 3 series and uses the stable API identifier gemini-3.6-flash.
Gemini 3.7 Flash is a proprietary, multimodal large language model released by Google on August 13, 2026.
Gemini 3.8 Flash is a multimodal model in Google's Gemini family. Google DeepMind released it on September 2, 2026 as a generally available model for software engineering, tool-using agents, and knowledge work.
Gemini Omni is a family of proprietary multimodal AI models from Google DeepMind for generating and editing media.
Gemini Robotics 2 is a family of three robotics models announced by Google DeepMind on July 30, 2026: a vision-language-action (VLA) model of the same name, an embodied reasoning model called Gemini Robotics…
Gemma 3 is a family of open-weight large language models developed by Google DeepMind and released on March 12, 2025.
GEN-1 is an embodied robot foundation model and control system developed by Generalist AI.
Generalist GEN-1.5 is a proprietary robot foundation model announced by Generalist AI on August 19, 2026.
ImageBind is a multimodal model from Meta AI (its Fundamental AI Research lab) that learns a single joint embedding space across six different modalities: images and video, text, audio, depth, thermal…
LINGO-2 is a closed-loop vision-language-action model for autonomous driving developed by the British self-driving company Wayve.
Llama 3.2 is a family of four open-weight large language models released by Meta on September 25, 2024, comprising lightweight 1 billion and 3 billion parameter text-only models for on-device AI and the 11…
Llama 4 Scout and Llama 4 Maverick are open-weight, natively multimodal AI large language models developed by Meta and released on April 5, 2025.
MAGI-2 Preview is a public research release of a unified audio-video generation model developed by Sand.ai.
Molmo is a family of open-weight, open-data vision-language models (VLMs) released by the Allen Institute for AI (Ai2) on 25 September 2024.
Muse Glimmer is an open-weight text-and-image model developed by Meta AI for local agent and coding workloads. Meta released the model on August 10, 2026 under the identifier meta-models/Muse-Glimmer-30B.
Muse Image is a proprietary image-generation and editing system developed by Meta Superintelligence Labs, a division of Meta.
Pika is an artificial intelligence video generation platform developed by Pika Labs, Inc. that lets users create and edit short videos from text prompts, images, and existing clips, and it is best known for…
Pixtral is a family of multimodal vision-language models developed by Mistral AI, a French AI company founded in April 2023.
Qwen3.8-Flash-Next is an experimental open-weight multimodal large language model released by Alibaba Group's Qwen team on August 26, 2026.
Reka Core is a frontier class multimodal foundation model developed by Reka AI, a research and product company founded in 2022 by former scientists from DeepMind, Google Brain, Meta FAIR, and Baidu.
Reka Edge is a 7-billion-parameter multimodal language model developed by Reka AI, introduced in April 2024 as the smallest member of the company's first publicly described model family.
Reka Flash is a family of multimodal large language models developed by Reka AI, a San Francisco Bay Area research company founded in 2022 by former researchers from Google DeepMind, Meta FAIR, and Google.
Robostral Navigate is an 8-billion-parameter vision language model developed by Mistral AI for instruction-following robot navigation.
SmolVLA (Small Vision-Language-Action) is a compact, open-source vision-language-action model (VLA) for robotics developed by Hugging Face and released in June 2025.
Sora 2 is a text-to-video and audio generation model developed by OpenAI, released on September 30, 2025, that OpenAI called "the GPT-3.5 moment for video." It succeeded the original Sora research preview from…
Text-to-image models are generative artificial intelligence systems that synthesize a new image from a natural-language description, called a prompt.
Visual question answering models are AI systems that take an image and a natural language question about that image and return a natural language answer.
WeMM-Embedding (WeChat Multi-Modal Embedding) is a family of open-weight universal multimodal embedding models built by the WeChat Vision team at Tencent.
dots3-note Preview is an open-weight multimodal model and large language model developed by dots studio, an AI team at Xiaohongshu.