ERQA
ERQA (Embodied Reasoning Question Answering) is a multimodal benchmark released by Google DeepMind in March 2025 to evaluate the embodied reasoning capabilities of vision-language models (VLMs) on robotics…
Explore Embodied AI through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of Embodied AI.
Showing 1-10 of 10 articles
ERQA (Embodied Reasoning Question Answering) is a multimodal benchmark released by Google DeepMind in March 2025 to evaluate the embodied reasoning capabilities of vision-language models (VLMs) on robotics…
Gemini Robotics 2 is a family of three robotics models announced by Google DeepMind on July 30, 2026: a vision-language-action (VLA) model of the same name, an embodied reasoning model called Gemini Robotics…
GEN-1 is an embodied robot foundation model and control system developed by Generalist AI.
Generalist GEN-1.5 is a proprietary robot foundation model announced by Generalist AI on August 19, 2026.
The original NVIDIA Cosmos Reason release is an open, customizable, 7-billion-parameter reasoning vision-language model (VLM) for physical AI and robotics developed by Nvidia.
PaLM-E (short for Pathways Language Model, Embodied) is an embodied multimodal large language model introduced by Google and TU Berlin in March 2023 that injects continuous sensor and image observations…
Robostral Navigate is an 8-billion-parameter vision language model developed by Mistral AI for instruction-following robot navigation.
SmolVLA (Small Vision-Language-Action) is a compact, open-source vision-language-action model (VLA) for robotics developed by Hugging Face and released in June 2025.
T-Rex: Tactile-Reactive Dexterous Manipulation is a 2026 robot learning research system for contact-rich, bimanual manipulation.
A vision-language-action model (VLA model, or VLA) is a machine-learning model for robotics that uses visual observations and a natural-language instruction to generate actions a robot can execute.