Multimodal AI

Explore Multimodal AI through related topics and the articles other pages reference most.

Explore articles

Reset filters
Browse subtopics: AI Agents

Articles that also belong to these categories. Counts cover all of Multimodal AI.

Showing 1-7 of 7 articles

CogAgent

CogAgent is an open visual language model built to act as a graphical user interface (GUI) agent: given a screenshot and a natural-language goal, it predicts the next on-screen action, such as where to click…

AI AgentsChinese AI

Computer-use agent

A computer-use agent (CUA) is a category of AI agent in artificial intelligence that performs tasks by directly operating a general-purpose computer's graphical user interface (GUI) the way a human does, by…

AI AgentsArtificial Intelligence

HuggingGPT

HuggingGPT is an agent system that uses a large language model as a controller to plan tasks and orchestrate specialist machine learning models hosted on Hugging Face.

AI AgentsAI Research

MM-BrowseComp

MM-BrowseComp is a benchmark for evaluating multimodal web-browsing AI agents, introduced in August 2025 by researchers from ByteDance, Nanjing University, M-A-P, the Institute of Automation of the Chinese…

AI AgentsAI Benchmarks

Muse Glimmer

Muse Glimmer is an open-weight text-and-image model developed by Meta AI for local agent and coding workloads. Meta released the model on August 10, 2026 under the identifier meta-models/Muse-Glimmer-30B.

AI AgentsAI Models

Project Astra

Project Astra is a research prototype from Google DeepMind that explores what a universal AI assistant might look like: a single agent that can see and hear the world in real time through a device camera and…

AI AgentsGoogle DeepMind

UI-TARS

UI-TARS is a native graphical user interface (GUI) agent model developed by ByteDance through its Seed research team.

AI AgentsChinese AI