Gemini Robotics 2
Gemini Robotics 2 is a family of three robotics models announced by Google DeepMind on July 30, 2026: a vision-language-action (VLA) model of the same name, an embodied reasoning model called Gemini Robotics…
Explore AI Models through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of AI Models.
Showing 1-17 of 17 articles
Gemini Robotics 2 is a family of three robotics models announced by Google DeepMind on July 30, 2026: a vision-language-action (VLA) model of the same name, an embodied reasoning model called Gemini Robotics…
GEN-1 is an embodied robot foundation model and control system developed by Generalist AI.
Generalist GEN-1.5 is a proprietary robot foundation model announced by Generalist AI on August 19, 2026.
Helix is a Vision-Language-Action (VLA) model for generalist humanoid control, developed in-house by Figure AI, the Sunnyvale, California humanoid-robot company founded by Brett Adcock.
Helix is a vision-language-action model (VLA) developed by Figure AI that controls humanoid robots by mapping camera images and natural language commands directly to continuous joint-level motion at 200 Hz.
Lumo-2 is a 4B-scale robot-learning model developed by Astribot. Astribot describes it as a latent world-action model: instead of rendering a future video before acting, it predicts an action-relevant…
Meta Motivo is a behavioral foundation model for controlling a simulated humanoid body, released by Meta AI's Fundamental AI Research (FAIR) group on December 12, 2024.
NVIDIA COMPASS is a framework and trained model for cross-embodiment robot navigation. Its name expands to Cross-Embodiment Mobility Policy via Residual RL and Skill Synthesis.
NVIDIA Cosmos is a world foundation model platform developed by NVIDIA for physical AI applications, including autonomous vehicles and robotics.
NVIDIA Isaac GR00T N1.7 is a 3-billion-parameter, cross-embodiment vision-language-action model developed by NVIDIA for humanoid and manipulation robots.
RFM-1 (Robotics Foundation Model 1) is an 8 billion parameter multimodal transformer for robotic manipulation announced by Covariant on March 11, 2024 at the MODEX 2024 trade show in Atlanta.
Robostral Navigate is an 8-billion-parameter vision language model developed by Mistral AI for instruction-following robot navigation.
Skild AI is a robotics artificial intelligence company building a general-purpose foundation model for physical embodiments
SmolVLA (Small Vision-Language-Action) is a compact, open-source vision-language-action model (VLA) for robotics developed by Hugging Face and released in June 2025.
A world action model (WAM) is a robot policy design that builds action generation on a video world model backbone rather than on a vision-language model, so that a single network jointly predicts how a scene…
π0.5 (also written pi0.5, pi 0.5, or π₀.₅, and pronounced "pi zero point five") is a vision-language-action model developed by the robotics company Physical Intelligence and released on April 22, 2025.
π₀ (pronounced pi-zero and sometimes written pi0 or pizero) is a vision-language-action model (VLA) developed by the robotics foundation-model startup Physical Intelligence