CogAgent
CogAgent is an open visual language model built to act as a graphical user interface (GUI) agent: given a screenshot and a natural-language goal, it predicts the next on-screen action, such as where to click…
Explore Multimodal AI through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of Multimodal AI.
Showing 1-7 of 7 articles
CogAgent is an open visual language model built to act as a graphical user interface (GUI) agent: given a screenshot and a natural-language goal, it predicts the next on-screen action, such as where to click…
A computer-use agent (CUA) is a category of AI agent in artificial intelligence that performs tasks by directly operating a general-purpose computer's graphical user interface (GUI) the way a human does, by…
HuggingGPT is an agent system that uses a large language model as a controller to plan tasks and orchestrate specialist machine learning models hosted on Hugging Face.
MM-BrowseComp is a benchmark for evaluating multimodal web-browsing AI agents, introduced in August 2025 by researchers from ByteDance, Nanjing University, M-A-P, the Institute of Automation of the Chinese…
Muse Glimmer is an open-weight text-and-image model developed by Meta AI for local agent and coding workloads. Meta released the model on August 10, 2026 under the identifier meta-models/Muse-Glimmer-30B.
Project Astra is a research prototype from Google DeepMind that explores what a universal AI assistant might look like: a single agent that can see and hear the world in real time through a device camera and…
UI-TARS is a native graphical user interface (GUI) agent model developed by ByteDance through its Seed research team.