Multimodal AI

Explore Multimodal AI through related topics and the articles other pages reference most.

Explore articles

Reset filters
Browse subtopics: Robotics

Articles that also belong to these categories. Counts cover all of Multimodal AI.

Showing 1-7 of 7 articles

Gemini Robotics 2

Gemini Robotics 2 is a family of three robotics models announced by Google DeepMind on July 30, 2026: a vision-language-action (VLA) model of the same name, an embodied reasoning model called Gemini Robotics…

AI ModelsEmbodied AI

SmolVLA

SmolVLA (Small Vision-Language-Action) is a compact, open-source vision-language-action model (VLA) for robotics developed by Hugging Face and released in June 2025.

AI HardwareAI Models

Vision-language-action model

A vision-language-action model (VLA model, or VLA) is a machine-learning model for robotics that uses visual observations and a natural-language instruction to generate actions a robot can execute.

Embodied AIRobotics