Gemini Robotics
Gemini Robotics is a family of robot foundation models developed by Google DeepMind that extends the Gemini multimodal model line into the physical world.
Explore AI Models through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of AI Models.
Showing 1-29 of 29 articles
Gemini Robotics is a family of robot foundation models developed by Google DeepMind that extends the Gemini multimodal model line into the physical world.
Gemini Robotics 2 is a family of three robotics models announced by Google DeepMind on July 30, 2026: a vision-language-action (VLA) model of the same name, an embodied reasoning model called Gemini Robotics…
GEN-1 is an embodied robot foundation model and control system developed by Generalist AI.
Generalist GEN-1.5 is a proprietary robot foundation model announced by Generalist AI on August 19, 2026.
Helix is a Vision-Language-Action (VLA) model for generalist humanoid control, developed in-house by Figure AI, the Sunnyvale, California humanoid-robot company founded by Brett Adcock.
Helix is a vision-language-action model (VLA) developed by Figure AI that controls humanoid robots by mapping camera images and natural language commands directly to continuous joint-level motion at 200 Hz.
Hydra-0 is an experimental world model for robot manipulation that represents actions as trajectories of visible points in an image.
Isaac GR00T N1.5 is an open vision-language-action model from NVIDIA built as a generalist foundation model for humanoid robots.
LingBot-VLA 2.0 is a 6-billion-parameter vision-language-action model developed by Robbyant, the embodied-intelligence unit of Ant Group.
Lumo-2 is a 4B-scale robot-learning model developed by Astribot. Astribot describes it as a latent world-action model: instead of rendering a future video before acting, it predicts an action-relevant…
NVIDIA COMPASS is a framework and trained model for cross-embodiment robot navigation. Its name expands to Cross-Embodiment Mobility Policy via Residual RL and Skill Synthesis.
NVIDIA Cosmos is a world foundation model platform developed by NVIDIA for physical AI applications, including autonomous vehicles and robotics.
NVIDIA Isaac GR00T N1 is an open foundation model for humanoid robots developed by NVIDIA and unveiled by Jensen Huang on March 18, 2025 at the company's annual GTC conference in San Jose, California.
NVIDIA Isaac GR00T N1.6 is a 3-billion-parameter vision-language-action model and cross-embodiment robot foundation model developed by Nvidia.
NVIDIA Isaac GR00T N1.7 is a 3-billion-parameter, cross-embodiment vision-language-action model developed by NVIDIA for humanoid and manipulation robots.
OpenPI (stylized openpi) is the open-source repository of robot foundation models, training code, and inference utilities published by Physical Intelligence, the San Francisco robotics and AI startup…
OpenVLA is a 7-billion-parameter open-source vision-language-action model (VLA) for robotic manipulation, released in June 2024 by a collaboration of researchers from Stanford University, UC Berkeley, the…
RFM-1 (Robotics Foundation Model 1) is an 8 billion parameter multimodal transformer for robotic manipulation announced by Covariant on March 11, 2024 at the MODEX 2024 trade show in Atlanta.
RT-2 (Robotic Transformer 2) is a vision-language-action model developed by Google DeepMind that enables robots to execute novel tasks by transferring knowledge from internet-scale vision-language pretraining…
Robostral Navigate is an 8-billion-parameter vision language model developed by Mistral AI for instruction-following robot navigation.
A robot foundation model is a large-scale machine learning model, typically based on the transformer architecture, that is pre-trained on broad, diverse datasets of robot interactions and then adapted to a…
Skild AI is a robotics artificial intelligence company building a general-purpose foundation model for physical embodiments
SmolVLA (Small Vision-Language-Action) is a compact, open-source vision-language-action model (VLA) for robotics developed by Hugging Face and released in June 2025.
V-JEPA 2 (Video Joint Embedding Predictive Architecture 2) is an open-source video world model released by Meta AI on June 11, 2025 that learns to understand, predict, and plan in the physical world by…
A world action model (WAM) is a robot policy design that builds action generation on a video world model backbone rather than on a vision-language model, so that a single network jointly predicts how a scene…
π*0.6 (written "Pi-star-0.6") is a vision-language-action robot foundation model developed by Physical Intelligence, a San Francisco robotics startup.
π0 (pronounced "pi-zero") is a vision-language-action model for general-purpose robot control developed by Physical Intelligence, a San Francisco-based robotics startup, and introduced on October 31, 2024.
π0.5 (also written pi0.5, pi 0.5, or π₀.₅, and pronounced "pi zero point five") is a vision-language-action model developed by the robotics company Physical Intelligence and released on April 22, 2025.
π₀ (pronounced pi-zero and sometimes written pi0 or pizero) is a vision-language-action model (VLA) developed by the robotics foundation-model startup Physical Intelligence