Gemma 3n
Gemma 3n is an open, mobile-first multimodal model from Google built to run locally on phones, tablets, laptops, and other resource-constrained hardware, accepting text, image, audio, and video as input and…
Explore Google DeepMind through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of Google DeepMind.
Showing 61-106 of 106 articles
Gemma 3n is an open, mobile-first multimodal model from Google built to run locally on phones, tablets, laptops, and other resource-constrained hardware, accepting text, image, audio, and video as input and…
Gemma 4 is a family of open-weight multimodal models developed by Google DeepMind and released on April 2, 2026.
GenCast is a probabilistic, ensemble-based machine-learning weather forecasting model developed by Google DeepMind that uses a diffusion model to generate large ensembles of possible 15-day weather…
Genie is an 11-billion-parameter generative interactive environment from Google DeepMind, described by its creators as the first foundation world model: it turns a single image, photo, or sketch into a…
Genie 2 is a foundation world model developed by Google DeepMind, unveiled on December 4, 2024.
Genie 3 is a general-purpose foundation world model developed by Google DeepMind and announced on August 5, 2025, that generates interactive, navigable 3D environments from a single text prompt and runs in…
Gopher is a 280-billion-parameter autoregressive transformer language model developed by DeepMind and described in a trio of companion papers released on December 8, 2021.
GraphCast is a machine-learning weather-forecasting model from Google DeepMind that produces a 10-day global forecast at 0.25-degree resolution in under one minute on a single Google Cloud TPU v4 device
Imagen is a family of text-to-image diffusion models developed by Google, first introduced in May 2022 and as of 2026 in its fourth generation (Imagen 4).
Imagen 2 is the second generation of Google's text-to-image diffusion model, developed by Google DeepMind and first announced for developers and enterprises on December 13, 2023.
Imagen 3 is a text-to-image generation model developed by Google DeepMind, announced at Google I/O on May 14, 2024 and progressively rolled out to users through mid-2024 and into 2025.
Imagen 4 is the fourth-generation text-to-image model developed by Google DeepMind, announced on May 20, 2025, at Google I/O 2025.
Ioannis Antonoglou is a Greek artificial intelligence researcher known as a co-creator of several of the landmark reinforcement learning systems built at Google DeepMind, including the Atari-playing Deep…
IsoDDE, short for Isomorphic Labs Drug Design Engine, is a unified computational drug design system developed by Isomorphic Labs, the Alphabet subsidiary spun out of Google DeepMind in 2021.
Isomorphic Labs is a London-based artificial intelligence company that designs small molecule and antibody drugs using deep learning models derived from the AlphaFold family of protein structure predictors.
John Jumper (born 1 January 1985) is an American computational chemist and biophysicist who led the development of AlphaFold, the artificial intelligence system that predicts the three-dimensional structure of…
Koray Kavukcuoglu is a Turkish computer scientist who leads Google DeepMind as its Senior Vice President, a role he assumed on 5 August 2026, and serves as Chief AI Architect of Google .
Lyria is a family of AI music generation models developed by Google DeepMind, spanning text-to-music synthesis, real-time interactive music performance, and full-length song composition.
Lyria 2 is a high-fidelity, text-to-music generation model built by Google DeepMind that turns text prompts into professional-grade instrumental audio.
Lyria 3.5 is a music generation model from Google DeepMind and the newest member of the Lyria family.
MuZero is a model-based reinforcement learning algorithm developed by DeepMind that masters Go, chess, shogi, and 57 Atari video games at superhuman or state of the art level without ever being told the rules…
Mustafa Suleyman CBE (born August 1984) is a British artificial intelligence entrepreneur who is Executive Vice President and CEO of Microsoft AI, the role he has held since March 19, 2024.
Nano Banana 2 is the public nickname for Gemini 3.1 Flash Image, an image generation and editing model released by Google DeepMind on 26 February 2026 .
Open X-Embodiment (OXE) is a large-scale collaborative robotics research initiative, led by Google DeepMind and announced in October 2023, that produced the largest open-source real robot dataset and a family…
Oriol Vinyals (born 1983, Sabadell, Catalonia, Spain) is a Spanish machine learning researcher who is a co-founder of Discovery Loop, a public benefit corporation launched on August 5, 2026 with the stated…
PaLM (Pathways Language Model) is a family of dense, decoder-only large language models developed by Google Research. Google announced the original family on April 4, 2022.
PaLM-E (short for Pathways Language Model, Embodied) is an embodied multimodal large language model introduced by Google and TU Berlin in March 2023 that injects continuous sensor and image observations…
Perceiver is a family of general-purpose neural network architectures from DeepMind built around attention and a small latent bottleneck.
Project Astra is a research prototype from Google DeepMind that explores what a universal AI assistant might look like: a single agent that can see and hear the world in real time through a device camera and…
Project Mariner was a research-prototype AI agent from Google DeepMind that browsed the web and took actions inside a user's Chrome browser, such as clicking links, filling forms, scrolling pages, and…
RT-2 (Robotic Transformer 2) is a vision-language-action model developed by Google DeepMind that enables robots to execute novel tasks by transferring knowledge from internet-scale vision-language pretraining…
RecurrentGemma is a family of open-weight language models released by Google DeepMind that is built on the Griffin architecture rather than the standard Transformer.
RoboCat is a self-improving foundation agent for robotic manipulation developed by Google DeepMind.
SIMA (Scalable Instructable Multiworld Agent) is a family of generalist embodied AI agents developed by Google DeepMind that follow free-form natural-language instructions to act in a wide range of…
Shane Legg (born 1973) is a New Zealand-born machine learning researcher and entrepreneur who co-founded the artificial intelligence laboratory DeepMind in 2010 with Demis Hassabis and Mustafa Suleyman
SigLIP (Sigmoid Loss for Language-Image Pre-training) is a family of vision-language encoders developed by researchers at Google DeepMind that pre-trains image and text encoders by treating each image-text…
SimpleQA Verified is a short-form factuality benchmark released by Google DeepMind and Google Research in September 2025 that measures the parametric knowledge of large language models using roughly 1,000…
SmolVLA (Small Vision-Language-Action) is a compact, open-source vision-language-action model (VLA) for robotics developed by Hugging Face and released in June 2025.
Sparrow is a research dialogue agent built by DeepMind and introduced on 22 September 2022.
SynthID is a family of digital watermarking technologies developed by Google DeepMind for marking and identifying content generated by generative AI systems.
TacticAI is an artificial intelligence system that gives football (soccer) coaches tactical advice on corner kicks.
Veo is a family of text-to-video generative AI models developed by Google DeepMind, and is best known as the first video model from a leading AI lab to natively generate synchronized audio (dialogue, sound…
Veo 2 is a text-to-video generative AI model developed by Google DeepMind and announced on December 16, 2024, the second major iteration of the Veo family.
Veo 3 is a video generation model developed by Google DeepMind and announced at Google I/O on May 20, 2025, and it is the first commercially available video generation model to natively produce synchronized…
Veo 3.1 is a video generation model released by Google DeepMind on October 15, 2025, as an incremental update to Veo 3.
WaveNet is a deep generative model for raw audio waveforms developed by DeepMind that synthesizes speech by predicting one waveform sample at a time, each conditioned on all the samples before it.