metir
metir
Docs
Download on App StoreGet it on Google PlayLog inSign up
Back to Blog
EmbeddingGemma
Google DeepMind
Embeddings
RAG
Open Source AI

EmbeddingGemma 2: Multimodal Open Embedding Model Explained

Google DeepMind's EmbeddingGemma 2 puts text, code, images, video and audio in one vector space under Apache 2.0. What that changes for search and RAG.

Metir AI TeamOctober 7, 20267 min read
EmbeddingGemma 2: Multimodal Open Embedding Model Explained

On October 6, 2026, Google DeepMind released EmbeddingGemma 2, an open multimodal embedding model that maps text, code, images, video and audio into a single vector space. The full model has 740 million parameters, is released under the Apache 2.0 licence, and is built on the architecture of Gemma 4. Embedding models rarely make headlines, but they sit underneath most retrieval-augmented generation (RAG) systems and semantic search, which makes a change in what they can represent, and under what licence, worth understanding. This article explains what an embedding model does, what is new in EmbeddingGemma 2, how it compares with its predecessor and with other open models, and what the licence change means.

Google logoGoogle
Gemma logoGemma
Gemini logoGemini
EmbeddingGemma 2 comes from Google DeepMind and is built on the Gemma 4 architecture.
740MParameters, full model270M for text and code only
+9.92MTEB Code gain68.76 to 78.68
8,192Token context4x the original
Apache 2.0LicenceReplaces the Gemma terms

What an embedding model does

An embedding model converts a piece of content into a list of numbers, called a vector, so that content with similar meaning ends up close together. A search system embeds every document once, stores the vectors, then embeds each query and returns the nearest stored vectors. In a RAG pipeline, those nearest items are the passages handed to a language model as context. The quality of the embedding therefore caps the quality of retrieval: a generator cannot use a passage the retriever never surfaced.

For years the practical constraint was that each modality needed its own model. Text embeddings could not be compared with image embeddings, so a team wanting to search screenshots, recordings and documents together ran separate indexes and stitched the results. A "single embedding space" removes that seam. Google states that EmbeddingGemma 2 places text, code, images, video and audio in one 768-dimensional space, and that inputs can be interleaved, according to the Hugging Face model card.

How single-space multimodal embeddings change RAG and search

The most direct effect is cross-modal retrieval. A text question can retrieve a video frame, a spoken passage or a diagram, because all of them live in the same coordinate system. A meeting recording no longer has to be transcribed before it becomes searchable, and a slide image no longer has to be captioned by a separate model. Skipping the intermediate transcript or caption also avoids losing whatever the transcriber missed.

The model card gives the cost of each input type in tokens against the 8,192-token context: 280 tokens per image by default, 140 per video frame, and 25 per second of audio. Video is sampled at one frame per second by default, and audio is processed as mono at 16 kHz. Google's launch material translates this into limits of about 5.5 minutes of audio, 29 images or 58 video frames per input, per the Google Developers Blog. Longer media must therefore be chunked, as long documents already are in text RAG.

“

A generator cannot use a passage the retriever never surfaced, so the embedding model sets the ceiling for RAG quality.

There are trade-offs worth keeping in view. A shared space is a compromise: one geometry has to serve very different content types, and a specialist single-modality model can still win on its home turf. Chunking choices, re-ranking and metadata filtering continue to matter. And multimodal indexes are larger and costlier to build than text-only ones, since a video generates many frames. The benchmark scores below are Google's own and should be read as claims to reproduce on your own data.

EmbeddingGemma 2 vs the original EmbeddingGemma

The original EmbeddingGemma, announced on September 4, 2025, was a text-only model. Google described it as 308 million parameters with a roughly 2,000-token context, and as the highest-ranking open multilingual text embedding model under 500 million parameters on MTEB at the time, per the original launch post. Its model card lists the licence as Gemma, which requires accepting Google's usage terms, and notes it was built on Gemma 3.

EmbeddingGemma 2 differs in four ways:

  • Modalities: it adds image, video and audio encoders (170 million and 300 million parameters) to a text backbone of 270 million parameters (130 million transformer plus 140 million embedder).
  • Context: 8,192 tokens against about 2,000.
  • Code retrieval: the MTEB code score rises from 68.76 to 78.68, the 9.92-point gain Google highlights. (The model card rounds the delta to +10.0.)
  • Multilingual text: essentially unchanged, at 61.36 against 61.15 on MTEB multilingual, so the text capability was kept rather than extended.

EmbeddingGemma 2 vs the original, at 768 dimensions

Text and code scores reported on the two Hugging Face model cards. Multilingual text is roughly flat; code retrieval carries the gain.

Sources: Hugging Face model cards for google/embeddinggemma-300m and google/embeddinggemma-2. Scores are Google-reported.

The new modalities are scored on benchmarks that have no v1 equivalent: 64.64 on MIEB for images, 50.67 Hit@1 on MMEB for video, and 69.54 MRR@10 on MSEB for audio, according to the model card. Without earlier baselines from the same family, those numbers are best compared against other models on the same benchmarks.

Matryoshka dimensions and on-device use

Both generations use Matryoshka Representation Learning, a training method that front-loads information into the first dimensions of a vector so it can be cut shorter with limited loss. EmbeddingGemma 2 outputs 768 dimensions natively and can be truncated to 512, 256 or 128. Google says slicing can cut storage by up to 8 times, and the model card states that 256 dimensions shows minimal quality loss while 128 degrades multimodal quality substantially. For a large index, the choice is a direct memory-versus-recall dial.

The model is also modular. Text and code need about 270 million parameters, adding vision brings it to 440 million, adding audio to 570 million, and the full model is 740 million. Every configuration shares one vector space, so an index built with the small configuration can in principle be queried from a device that loads a different one.

One checkpoint, four loadable sizes

Parameters (millions) when only the encoders a use case needs are loaded. Every configuration shares the same vector space.

Sources: Hugging Face model card for google/embeddinggemma-2; Google Developers Blog.

Google reports, for a Google Pixel 11 Pro with quantization, roughly 191 MB of active RAM for the text-only weights and about 567 MB for the full multimodal model, using INT4 and INT8 quantization-aware training and the LiteRT runtime. It also cites 37.3 ms per image embedding (26.9 images per second) on a MacBook M5 Pro GPU. These are vendor measurements on specific hardware. The practical implication is that private, offline semantic search over a user's own photos, audio and notes becomes feasible without sending content to a server.

The Google sign at 1600 Amphitheatre Parkway in Mountain View, California
The Google sign at 1600 Amphitheatre, Mountain View, California. File photo of Google's headquarters, not related to the release itself. Photo: Hakan Dahlstrom, CC BY 2.0.

How it compares with other open embedding models

The open field is crowded. Alibaba's Qwen3-Embedding-0.6B, released June 5, 2025 under Apache 2.0, has 0.6 billion parameters, a 32,000-token context, output dimensions from 32 to 1,024, and reports a 64.33 mean on MTEB multilingual, per its model card. That is higher than EmbeddingGemma 2's 61.36, and its context is far longer, although its model card describes a text embedding model, with no multimodal inputs listed. Direct comparison is also imperfect: scores come from each vendor's own evaluation, and benchmark versions and settings can differ.

So the claim in EmbeddingGemma 2 is less "best text embedder" than "one small model that spans five modalities with text quality held at the previous level." Google frames it as leading among sub-1-billion-parameter multimodal models, including on the Massive Audio Embedding Benchmark, which independent evaluations will need to confirm.

Why the Apache 2.0 licence matters

The original EmbeddingGemma shipped under the Gemma terms, a custom licence that users must accept and that carries Google's usage policy. EmbeddingGemma 2 is Apache 2.0, a standard permissive licence that most legal teams already have a position on. For enterprises, that can remove a review step, simplify redistribution inside products, and make the model easier to bundle in on-device apps. The weights are on Hugging Face and Kaggle, with support listed for transformers.js, llama.cpp, Ollama, vLLM, MLX, SGLang and LM Studio. Why Google changed the terms is not stated in the sources we reviewed. For wider context on the Gemma family's reach, see our post on Gemma passing 1 billion downloads.

What to watch next

  • Independent benchmarks: third-party reproductions on code, video and audio retrieval, and on real corpora rather than leaderboards.
  • Index economics: how much a multimodal index costs in storage and compute at 768 versus 256 dimensions.
  • Ecosystem uptake: vector databases and RAG frameworks adding native support for mixed-modality collections.
  • Rivals' responses: whether other labs release open multimodal embedders under permissive licences.

For teams building search and RAG, the practical question is whether one shared space beats several specialist indexes on their own data. Workspaces that route across many models, such as Metir, tend to benefit when retrieval components are open and swappable, but the answer will come from testing, not from launch-day tables.

Sources:

  • Google Developers Blog: Bring multimodal semantic search to the edge with EmbeddingGemma 2
  • Google blog: EmbeddingGemma 2 is a best-in-class open model for natively multimodal embeddings
  • Hugging Face: google/embeddinggemma-2
  • Hugging Face: google/embeddinggemma-300m
  • Google Developers Blog: Introducing EmbeddingGemma (Sept 4, 2025)
  • Hugging Face: Qwen/Qwen3-Embedding-0.6B

Image credits

  • Header image: entrance to the Google building at 6 Pancras Square, King's Cross, London, which houses Google DeepMind, by Gciriani via Wikimedia Commons, licensed under CC BY-SA 4.0.
  • In-body image: Google sign at 1600 Amphitheatre, Mountain View, by Hakan Dahlstrom via Wikimedia Commons, licensed under CC BY 2.0.

Ready to experience AI that adapts to you?

metir brings together the world's best AI models in one seamless experience. Start for free today.

Get Started Free
metir

Agentic Operating System for Professionals buried in meetings, emails and docs.

© 2026 metir. All rights reserved.

Product

  • Features
  • Pricing
  • Research
  • Docs
  • Blog
  • Enterprise

Company

  • Docs
  • Support
  • Careers

Legal

  • Terms of service
  • Privacy policy

Personalisation is powerful. Privacy is non-negotiable.

Status: All systems operational