Citation and evidence

EmbeddingGemma 2

9 min full readUpdated 12 references

This article's verification

Report a problem with this article

More

Use this article

Raw MarkdownExplore connections

Improve this page

Suggest editRevision historyDiscussion

Browse categories

Information RetrievalOpen Source AISmall Language Models

Cite this article

EmbeddingGemma 2 is an open multimodal embedding model released by Google DeepMind on October 6, 2026. It converts text, source code, images, video and audio into vectors that can be compared in a shared space. The model is distributed under the Apache 2.0 license.[1]

Built on Gemma 4, it has approximately 740 million parameters with all modality encoders loaded. Its default output is a 768-dimensional vector, with shorter representations available through Matryoshka Representation Learning. Google documents support for more than 100 languages and an 8,192-token input context.[2]

EmbeddingGemma 2 is distinct from the original EmbeddingGemma, a text embedding model based on Gemma 3. The second version adds native media inputs and expands the documented context from 2,048 to 8,192 tokens.[2]

Purpose and retrieval workflows

The model produces numerical representations rather than conversational answers. A text query can be compared with an image, sound recording or video representation without first converting that media into a caption or transcript. Text and media can also be embedded together as one item.[1]

In retrieval-augmented generation, it retrieves relevant documents. Google's guide also demonstrates zero-shot classification by comparing inputs with candidate labels, without task-specific training.[6]

Architecture

The released configuration identifies an EmbeddingGemma2Model with a text backbone and separate vision and audio components. Its text stack alternates five local-attention layers with one full-attention layer. The configuration lists two key-value heads for local layers and one for the full-attention layers.[3]

ComponentSelected configuration values
Text stack24 layers; hidden size 512; feed-forward size 2,048
Text attentionFour attention heads; five local layers per full-attention layer
Vocabulary262,144 entries
Embedding output768 dimensions
Vision encoder16 layers; hidden size 768; 12 attention heads; patch size 16
Audio encoder12 layers; hidden size 1,024; eight attention heads
Published weight dtypebfloat16

Expanded article table

These values describe the released configuration file, not measured latency or memory requirements. The supported inference context is the 8,192 tokens stated in Google's model documentation; larger positional-capacity fields in a configuration file should not be substituted for that documented input limit.[2][3]

Selectively loading encoders

The checkpoint supports four encoder configurations with compatible output vectors.[4]

Inputs requiredApproximate parametersSentence Transformers configuration override
Text and code270 million{"vision_config": None, "audio_config": None}
Text, images and video frames440 million{"audio_config": None}
Text and audio570 million{"vision_config": None}
All modalities740 millionNo override

Expanded article table

Within the same checkpoint and output configuration, encoder selection preserves query/document compatibility without re-embedding.[4]

The parameter totals are rounded. Google's multimodal tutorial reports 744,371,488 parameters for its full-model example, rather than exactly 740 million.[5]

Multimodal inputs and context

The Sentence Transformers interface accepts media dictionaries such as {"image": source}, {"audio": source} and {"video": source}. Sources can include files, URLs and supported in-memory objects. A dictionary with several modality keys combines those inputs; <|image|>, <|audio|> and <|video|> placeholders in its text can specify their order.[5]

All modalities consume the same context budget. The card gives these default token costs.[8]

InputDefault token cost
Image280 per image
Video140 per sampled frame
Audio25 per second

Expanded article table

The launch announcement describes capacities of about 29 images, 58 video frames or 5.5 minutes of audio within the 8K context. These are alternatives, not simultaneous allowances for an interleaved input.[1]

The multimodal guide documents configurable vision budgets of 70, 140, 280, 560 or 1,120 tokens. Larger budgets consume more context and processing time. Video is sampled at one frame per second by default, with configurable frame limits and sampling behavior. Audio is decoded and resampled to 16 kHz.[5]

Frame sampling matters when interpreting video search: the interface processes selected frames rather than guaranteeing inspection of every frame in a recording.[5]

Text prompts and a minimal example

Retrieval uses different query and document prefixes; symmetric tasks use matching prefixes.[6]

TaskPrompt name
Search querySearchQuery
Untitled documentDocument
Question-to-passage retrievalQuestionAnswering
Claim-to-evidence retrievalFactChecking
Code retrievalCodeRetrieval
ClassificationClassification
ClusteringClustering
Sentence similaritySentenceSimilarity or STS

Expanded article table

The Document prompt sets a placeholder title. Real titles can be included in a custom document prefix. Text-task prefixes are not applied to media.[6]

Google's developer guide specifies Sentence Transformers 6.1.0 or later for this model.[4] A text-only retrieval call can be written as follows:[6]

from sentence_transformers import SentenceTransformer

embedder = SentenceTransformer(
    "google/embeddinggemma-2",
    config_kwargs={"vision_config": None, "audio_config": None},
)
question_vec = embedder.encode(
    "Where is the deployment checklist?", prompt_name="SearchQuery"
)
passage_vec = embedder.encode(
    "The release notes include a deployment checklist.", prompt_name="Document"
)
scores = embedder.similarity(question_vec, passage_vec)

Google's multimodal guide warns that the model's activations are incompatible with float16. It recommends bfloat16 on supported hardware or float32.[5]

Shorter vectors and storage

Matryoshka Representation Learning trains nested prefixes of a representation to remain useful at different lengths. The original academic method optimizes losses at several representation sizes within one vector; it is not ordinary slicing of an arbitrary embedding model.[7]

Supported dimensions are 768, 512, 256 and 128. Use truncate_dim with normalize_embeddings=True; query and index dimensions must match.[4]

At equal precision, raw vectors require one-third the storage at 256 dimensions, or one-sixth at 128. This excludes documents, index overhead and model weights.[4][7]

The card's overall MMEB v2 score is 59.01 at 768 dimensions, 56.24 at 256 and 45.65 at 128.[8]

Reported evaluation results

Google reports the following results for the full-precision checkpoint at 768 dimensions.[8]

BenchmarkAggregation and metricScore
MTEB multilingual v2Task mean; mixed metrics61.36
MTEB code v1Task mean; NDCG@1078.68
MIEB liteTask-type mean; mixed metrics64.64
MMEB v2 imageTask mean; Hit@157.28
MMEB v2 visual documentsTask mean; NDCG@567.84
MMEB v2 videoTask mean; Hit@150.67
MSEB retrievalTask mean; MRR@1069.54
MAEBTask mean; mixed metrics49.39

Expanded article table

The reported code score rises from the original model's 68.76 to 78.68: a 9.92-point gain.[8]

What the benchmarks measure

The original MTEB paper evaluates several embedding uses rather than one universal accuracy measure. Its retrieval tasks rank relevant documents and use NDCG@10 as the main metric, while classification, clustering and semantic similarity use other measures. The authors found that models' strengths varied across tasks.[9]

MMTEB extends that evaluation framework across more languages and domains. Its methodology distinguishes an average across tasks from an average across task categories and from Borda-count ranking. Consequently, a table labeled as a task mean should not be read as a rank or silently changed to a category-weighted average.[10]

MAEB evaluates audio representations across speech, music, environmental sound and audio-text tasks. Its paper explicitly excludes transcription and generation from the capabilities it measures. The benchmark's aggregate score therefore does not establish speech-recognition accuracy or audio-generation quality.[11]

The rows above retain Google's stated benchmark versions and metrics. They are not a single common percentage of correctness, and they should not be compared across rows as if each measured the same task.[9][11]

Local deployment and availability

At launch, Google published weights on Hugging Face and Kaggle and listed integrations with tools including Transformers, Sentence Transformers, MLX, vLLM, llama.cpp, SGLang, Ollama and LM Studio. Its announcement described Gemini Enterprise Agent Platform Model Garden availability as forthcoming, not already available.[1]

Google's edge announcement reports approximately 191 MB of active RAM for text-only weights and 567 MB for the full model on a Pixel 11 Pro. These are device-specific figures associated with its optimized edge implementation, not a promise that an unquantized Python process uses the same memory.[12]

The same announcement reports 37.3 ms per visual embedding on a MacBook M5 Pro GPU with a maximum budget of 70 vision tokens per image. That token budget is lower than the default image allocation above, so the figure should not be presented as default-resolution performance on arbitrary hardware.[12]

The edge release includes Instant Media Search and Video Moments Finder demonstrations in Google AI Edge Gallery. It describes MediaPipe embedding, retrieval and decision tasks and LiteRT deployment bundles. ML Kit integration was announced for the following weeks, rather than as an available launch-day Android service.[12]

Limitations and application safeguards

The card warns of unequal language performance and says the model lacks post-training alignment and output moderation. Applications need retrieval filtering and fairness evaluation.[8]

Google's text tutorial gives unrelated sentences a relatively high cosine similarity. A similarity score is not a probability of truth.[6]

References

  1. ^1 ^2 ^3 ^4Dua, Sahil, and Henrique Schechter Vera. "EmbeddingGemma 2: an open, lightweight multimodal embedding model." Google, October 6, 2026. Official launch article.
  2. ^1 ^2 ^3Google AI for Developers. "EmbeddingGemma overview." Updated October 6, 2026. Model overview and previous versions.
  3. ^1 ^2Google. "EmbeddingGemma 2 configuration." Hugging Face model repository, accessed October 7, 2026. Released config.json.
  4. ^1 ^2 ^3 ^4 ^5Grootendorst, Maarten, and Ian Ballantyne. "EmbeddingGemma 2: The Developer Guide." Google Developers Blog, October 6, 2026. Developer guide.
  5. ^1 ^2 ^3 ^4 ^5Google AI for Developers. "Multimodal Embeddings with EmbeddingGemma 2." Updated October 6, 2026. Multimodal interface guide.
  6. ^1 ^2 ^3 ^4 ^5Google AI for Developers. "Generating Embeddings with EmbeddingGemma 2 and Sentence Transformers." Updated October 6, 2026. Text inference guide.
  7. ^1 ^2Kusupati, Aditya, et al. "Matryoshka Representation Learning." arXiv:2205.13147, 2022, version 4, February 8, 2024. Paper.
  8. ^1 ^2 ^3 ^4 ^5Google DeepMind. "EmbeddingGemma 2 model card." Accessed October 7, 2026. Official model card, Hugging Face mirror.
  9. ^1 ^2Muennighoff, Niklas, et al. "MTEB: Massive Text Embedding Benchmark." arXiv:2210.07316, 2022, version 3, March 19, 2023. Paper.
  10. ^Enevoldsen, Kenneth, et al. "MMTEB: Massive Multilingual Text Embedding Benchmark." arXiv:2502.13595, 2025, version 4, November 13, 2025. Paper.
  11. ^1 ^2El Assadi, Adnan, et al. "MAEB: Massive Audio Embedding Benchmark." arXiv:2602.16008, February 17, 2026. Paper.
  12. ^1 ^2 ^3Google AI Edge Team and Android ML Team. "Bring multimodal semantic search to the edge with EmbeddingGemma 2." Google Developers Blog, October 6, 2026. Edge deployment article.

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

v1 · 1,734 words · full history

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here

Reviewer note: Independent full-article review against 12 primary and research references, October 7, 2026. Checked subject identity, specifications, availability, limitations and citation support.

Cite this page: AI Wiki. "EmbeddingGemma 2." aiwiki.ai, updated 6 Oct 2026, fact-checked 6 Oct 2026. CC BY 4.0. https://aiwiki.ai/wiki/embeddinggemma_2

Suggest edit