Search across media with the same query

EmbeddingGemma 2 is a 740-million-parameter model for mapping text, images, video and audio into one embedding space. Embeddings turn content into vectors that software can compare for similarity. Google’s example is practical: find a particular video clip using a voice memo, or search hours of recordings with a text query, using one multimodal model.

Google is releasing it under Apache 2.0 and designing it for on-device inference. That gives developers a way to build retrieval systems without sending every clip or recording to a hosted model, depending on their app and device setup.

A modular model for smaller devices

The model is modular: Google describes a 270-million-parameter text-only configuration, with optional vision and audio encoders for broader media support. Its Matryoshka representation learning approach allows output vectors to be shortened from 768 dimensions to 128, which the company says can reduce local vector storage and memory needs by up to six times.

On a Pixel 11 Pro, Google reports that quantized text-only weights can use around 191 MB of active RAM and the full multimodal model around 567 MB. Those figures are Google’s measurements on a particular device; actual performance will depend on quantization, device, workload and app design.

What developers can build

The shared vector space makes cross-modal retrieval more direct: a text question can match a moment in video, a spoken note can locate a related clip, and a transcript can be connected to images or code. It is a retrieval component, not a generative assistant that understands a whole archive automatically.

For builders, the best fit is an app with private media or many recordings: meetings, field notes, personal video libraries or support archives. The new option is to move more of the indexing and search process onto the user’s device, then compare quality and memory use against the existing cloud pipeline.

Explore the original source ↗

Source published October 6, 2026. Coverage is based on the maker’s announcement and demonstration.