All News
googlegemmaembeddingson-devicemultimodal

Google DeepMind's EmbeddingGemma 2 puts one vector space for text, images and audio on the phone

Google DeepMind's EmbeddingGemma 2 fits one open 740M-parameter vector space for text, images, video and audio on your phone, with 6x smaller embeddings.

Vlad MakarovVlad Makarovreviewed and published
2 min read
Google DeepMind's EmbeddingGemma 2 puts one vector space for text, images and audio on the phone

Google DeepMind launched EmbeddingGemma 2 on October 6, an open-weight embedding model under a commercially permissive Apache 2.0 licence. It has 740 million parameters and maps text, including code, images, video frames and audio into one unified vector space, all running on a phone. The benchmark gains are real: MTEB Code rose from 68.76 to 78.68. But the quieter number is the vector itself. Native output is 768 dimensions, truncatable to as few as 128, for up to six times less storage. For on-device retrieval, that is the figure that decides what actually ships.

What Happened

EmbeddingGemma 2 is built on the Gemma 4 architecture and, per its model card, shares Gemma 4's text tokenizer and audio encoder, so the two models can run in one pipeline with a lower combined memory footprint. The encoders are modular, and developers load only the modalities they need:

  • Total size: 740M parameters (270M text, 170M vision, 300M audio)
  • Vectors: 768 dimensions, truncatable to 512, 256 or 128
  • Context: 8K tokens, or 5.5 minutes of audio, 29 images, 58 video frames
  • Footprint: about 191MB active RAM for text-only weights, about 567MB fully multimodal, on a Pixel 11 Pro
  • Code: MTEB Code 78.68, a 9.92-point gain over EmbeddingGemma 1

Google says the model also works as a zero-shot intent-routing engine, matching inputs against classification labels in milliseconds with no training data or fine-tuning. The first EmbeddingGemma passed 20 million downloads.

Why This Matters

One vector space for text, images, video and audio is the headline. The storage arithmetic is the story. Every indexed file, frame or recording leaves a vector behind, and at 768 dimensions those vectors add up fast on a device with a fixed flash budget. Matryoshka Representation Learning lets developers keep only the leading dimensions, and the model card reports quality holding close to lossless down to 256 dimensions, with 128 reserved for text-only work. That is a six-fold cut in vector-database size for a small, measurable quality cost, the difference between an on-device index that fits and one that does not. It is the same constraint driving local inference more broadly, as our look at the iPhone's second GPU argued.

What's Next

Google says EmbeddingGemma 2 will arrive as a service on Android through ML Kit in the coming weeks, with NPU acceleration for supported devices, and it already ships in the AI Edge Gallery's Instant Media Search and Video Moments Finder. The Hacker News thread drew more than 400 points. The open question is whether 128-dimensional text embeddings hold up against a specific workload, or only on aggregate.

Related Articles

Scroll down

to load the next article