EmbeddingGemma 2
EmbeddingGemma 2 is an open multimodal embedding model released by Google DeepMind on October 6, 2026. It converts text, source code, images, video and audio into vectors that can be compared in a shared space. The model is distributed under the Apache 2.0 license.[1]
Built on Gemma 4, it has approximately 740 million parameters with all modality encoders loaded. Its default output is a 768-dimensional vector, with shorter representations available through Matryoshka Representation Learning. Google documents support for more than 100 languages and an 8,192-token input context.[2]
EmbeddingGemma 2 is distinct from the original EmbeddingGemma, a text embedding model based on Gemma 3. The second version adds native media inputs and expands the documented context from 2,048 to 8,192 tokens.[2]
Purpose and retrieval workflows
The model produces numerical representations rather than conversational answers. A text query can be compared with an image, sound recording or video representation without first converting that media into a caption or transcript. Text and media can also be embedded together as one item.[1]
In retrieval-augmented generation, it retrieves relevant documents. Google's guide also demonstrates zero-shot classification by comparing inputs with candidate labels, without task-specific training.[6]
Architecture
The released configuration identifies an EmbeddingGemma2Model with a text backbone and separate vision and audio components. Its text stack alternates five local-attention layers with one full-attention layer. The configuration lists two key-value heads for local layers and one for the full-attention layers.[3]
| Component | Selected configuration values |
|---|---|
| Text stack | 24 layers; hidden size 512; feed-forward size 2,048 |
| Text attention | Four attention heads; five local layers per full-attention layer |
| Vocabulary | 262,144 entries |
| Embedding output | 768 dimensions |
| Vision encoder | 16 layers; hidden size 768; 12 attention heads; patch size 16 |
| Audio encoder | 12 layers; hidden size 1,024; eight attention heads |
| Published weight dtype | bfloat16 |
These values describe the released configuration file, not measured latency or memory requirements. The supported inference context is the 8,192 tokens stated in Google's model documentation; larger positional-capacity fields in a configuration file should not be substituted for that documented input limit.[2][3]
Selectively loading encoders
The checkpoint supports four encoder configurations with compatible output vectors.[4]
| Inputs required | Approximate parameters | Sentence Transformers configuration override |
|---|---|---|
| Text and code | 270 million | {"vision_config": None, "audio_config": None} |
| Text, images and video frames | 440 million | {"audio_config": None} |
| Text and audio | 570 million | {"vision_config": None} |
| All modalities | 740 million | No override |
Within the same checkpoint and output configuration, encoder selection preserves query/document compatibility without re-embedding.[4]
The parameter totals are rounded. Google's multimodal tutorial reports 744,371,488 parameters for its full-model example, rather than exactly 740 million.[5]
Multimodal inputs and context
The Sentence Transformers interface accepts media dictionaries such as {"image": source}, {"audio": source} and {"video": source}. Sources can include files, URLs and supported in-memory objects. A dictionary with several modality keys combines those inputs; <|image|>, <|audio|> and <|video|> placeholders in its text can specify their order.[5]
All modalities consume the same context budget. The card gives these default token costs.[8]
| Input | Default token cost |
|---|---|
| Image | 280 per image |
| Video | 140 per sampled frame |
| Audio | 25 per second |
The launch announcement describes capacities of about 29 images, 58 video frames or 5.5 minutes of audio within the 8K context. These are alternatives, not simultaneous allowances for an interleaved input.[1]
The multimodal guide documents configurable vision budgets of 70, 140, 280, 560 or 1,120 tokens. Larger budgets consume more context and processing time. Video is sampled at one frame per second by default, with configurable frame limits and sampling behavior. Audio is decoded and resampled to 16 kHz.[5]
Frame sampling matters when interpreting video search: the interface processes selected frames rather than guaranteeing inspection of every frame in a recording.[5]
Text prompts and a minimal example
Retrieval uses different query and document prefixes; symmetric tasks use matching prefixes.[6]
| Task | Prompt name |
|---|---|
| Search query | SearchQuery |
| Untitled document | Document |
| Question-to-passage retrieval | QuestionAnswering |
| Claim-to-evidence retrieval | FactChecking |
| Code retrieval | CodeRetrieval |
| Classification | Classification |
| Clustering | Clustering |
| Sentence similarity | SentenceSimilarity or STS |
The Document prompt sets a placeholder title. Real titles can be included in a custom document prefix. Text-task prefixes are not applied to media.[6]
Google's developer guide specifies Sentence Transformers 6.1.0 or later for this model.[4] A text-only retrieval call can be written as follows:[6]
from sentence_transformers import SentenceTransformer
embedder = SentenceTransformer(
"google/embeddinggemma-2",
config_kwargs={"vision_config": None, "audio_config": None},
)
question_vec = embedder.encode(
"Where is the deployment checklist?", prompt_name="SearchQuery"
)
passage_vec = embedder.encode(
"The release notes include a deployment checklist.", prompt_name="Document"
)
scores = embedder.similarity(question_vec, passage_vec)
Google's multimodal guide warns that the model's activations are incompatible with float16. It recommends bfloat16 on supported hardware or float32.[5]
Shorter vectors and storage
Matryoshka Representation Learning trains nested prefixes of a representation to remain useful at different lengths. The original academic method optimizes losses at several representation sizes within one vector; it is not ordinary slicing of an arbitrary embedding model.[7]
Supported dimensions are 768, 512, 256 and 128. Use truncate_dim with normalize_embeddings=True; query and index dimensions must match.[4]
At equal precision, raw vectors require one-third the storage at 256 dimensions, or one-sixth at 128. This excludes documents, index overhead and model weights.[4][7]
The card's overall MMEB v2 score is 59.01 at 768 dimensions, 56.24 at 256 and 45.65 at 128.[8]
Reported evaluation results
Google reports the following results for the full-precision checkpoint at 768 dimensions.[8]
| Benchmark | Aggregation and metric | Score |
|---|---|---|
| MTEB multilingual v2 | Task mean; mixed metrics | 61.36 |
| MTEB code v1 | Task mean; NDCG@10 | 78.68 |
| MIEB lite | Task-type mean; mixed metrics | 64.64 |
| MMEB v2 image | Task mean; Hit@1 | 57.28 |
| MMEB v2 visual documents | Task mean; NDCG@5 | 67.84 |
| MMEB v2 video | Task mean; Hit@1 | 50.67 |
| MSEB retrieval | Task mean; MRR@10 | 69.54 |
| MAEB | Task mean; mixed metrics | 49.39 |
The reported code score rises from the original model's 68.76 to 78.68: a 9.92-point gain.[8]
What the benchmarks measure
The original MTEB paper evaluates several embedding uses rather than one universal accuracy measure. Its retrieval tasks rank relevant documents and use NDCG@10 as the main metric, while classification, clustering and semantic similarity use other measures. The authors found that models' strengths varied across tasks.[9]
MMTEB extends that evaluation framework across more languages and domains. Its methodology distinguishes an average across tasks from an average across task categories and from Borda-count ranking. Consequently, a table labeled as a task mean should not be read as a rank or silently changed to a category-weighted average.[10]
MAEB evaluates audio representations across speech, music, environmental sound and audio-text tasks. Its paper explicitly excludes transcription and generation from the capabilities it measures. The benchmark's aggregate score therefore does not establish speech-recognition accuracy or audio-generation quality.[11]
The rows above retain Google's stated benchmark versions and metrics. They are not a single common percentage of correctness, and they should not be compared across rows as if each measured the same task.[9][11]
Local deployment and availability
At launch, Google published weights on Hugging Face and Kaggle and listed integrations with tools including Transformers, Sentence Transformers, MLX, vLLM, llama.cpp, SGLang, Ollama and LM Studio. Its announcement described Gemini Enterprise Agent Platform Model Garden availability as forthcoming, not already available.[1]
Google's edge announcement reports approximately 191 MB of active RAM for text-only weights and 567 MB for the full model on a Pixel 11 Pro. These are device-specific figures associated with its optimized edge implementation, not a promise that an unquantized Python process uses the same memory.[12]
The same announcement reports 37.3 ms per visual embedding on a MacBook M5 Pro GPU with a maximum budget of 70 vision tokens per image. That token budget is lower than the default image allocation above, so the figure should not be presented as default-resolution performance on arbitrary hardware.[12]
The edge release includes Instant Media Search and Video Moments Finder demonstrations in Google AI Edge Gallery. It describes MediaPipe embedding, retrieval and decision tasks and LiteRT deployment bundles. ML Kit integration was announced for the following weeks, rather than as an available launch-day Android service.[12]
Limitations and application safeguards
The card warns of unequal language performance and says the model lacks post-training alignment and output moderation. Applications need retrieval filtering and fairness evaluation.[8]
Google's text tutorial gives unrelated sentences a relatively high cosine similarity. A similarity score is not a probability of truth.[6]
References
- ^1 ^2 ^3 ^4Dua, Sahil, and Henrique Schechter Vera. "EmbeddingGemma 2: an open, lightweight multimodal embedding model." Google, October 6, 2026. Official launch article.
- ^1 ^2 ^3Google AI for Developers. "EmbeddingGemma overview." Updated October 6, 2026. Model overview and previous versions.
- ^1 ^2Google. "EmbeddingGemma 2 configuration." Hugging Face model repository, accessed October 7, 2026. Released config.json.
- ^1 ^2 ^3 ^4 ^5Grootendorst, Maarten, and Ian Ballantyne. "EmbeddingGemma 2: The Developer Guide." Google Developers Blog, October 6, 2026. Developer guide.
- ^1 ^2 ^3 ^4 ^5Google AI for Developers. "Multimodal Embeddings with EmbeddingGemma 2." Updated October 6, 2026. Multimodal interface guide.
- ^1 ^2 ^3 ^4 ^5Google AI for Developers. "Generating Embeddings with EmbeddingGemma 2 and Sentence Transformers." Updated October 6, 2026. Text inference guide.
- ^1 ^2Kusupati, Aditya, et al. "Matryoshka Representation Learning." arXiv:2205.13147, 2022, version 4, February 8, 2024. Paper.
- ^1 ^2 ^3 ^4 ^5Google DeepMind. "EmbeddingGemma 2 model card." Accessed October 7, 2026. Official model card, Hugging Face mirror.
- ^1 ^2Muennighoff, Niklas, et al. "MTEB: Massive Text Embedding Benchmark." arXiv:2210.07316, 2022, version 3, March 19, 2023. Paper.
- ^Enevoldsen, Kenneth, et al. "MMTEB: Massive Multilingual Text Embedding Benchmark." arXiv:2502.13595, 2025, version 4, November 13, 2025. Paper.
- ^1 ^2El Assadi, Adnan, et al. "MAEB: Massive Audio Embedding Benchmark." arXiv:2602.16008, February 17, 2026. Paper.
- ^1 ^2 ^3Google AI Edge Team and Android ML Team. "Bring multimodal semantic search to the edge with EmbeddingGemma 2." Google Developers Blog, October 6, 2026. Edge deployment article.
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
v1 · 1,734 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: Independent full-article review against 12 primary and research references, October 7, 2026. Checked subject identity, specifications, availability, limitations and citation support.
Cite this page: AI Wiki. "EmbeddingGemma 2." aiwiki.ai, updated 6 Oct 2026, fact-checked 6 Oct 2026. CC BY 4.0. https://aiwiki.ai/wiki/embeddinggemma_2