EmbeddingGemma 2: Google DeepMind's 740M open model brings multimodal search to your phone, fully offline
Google DeepMind launched EmbeddingGemma 2, a 740-million-parameter open model that maps text, code, images, video and audio into one shared embedding space — small enough to run entirely on-device under the Apache 2.0 license.
On Tuesday, Google DeepMind launched EmbeddingGemma 2 — a 740-million-parameter embedding model that natively maps text, code, images, video and audio into one shared embedding space. Built on the Gemma 4 architecture and released under the commercially permissive Apache 2.0 license, it's designed for inference on consumer hardware: on a Pixel 11 Pro, the text-only version uses about 191MB of active RAM, the full multimodal model about 567MB.
The pitch is simple: find a specific video clip from a voice memo, or search hours of recorded audio with a text query, and do it all without your data ever leaving the device.
What changed from version one
EmbeddingGemma launched last year as a text-only embedding model for on-device search, and Google says developer response beat its expectations: more than 20 million downloads, powering smarter on-device search tools and privacy-first retrieval-augmented generation pipelines. EmbeddingGemma 2 extends that same playbook beyond text.
The new model is modular by design. Text-only workloads need just 270M parameters, with optional vision (170M) and audio (300M) encoders bolted on for full multimodal support. Context is up 4x to 8K tokens, enough to process up to 5.5 minutes of audio, 29 images, 58 video frames, or interleaved combinations — all locally.
Storage is handled with Matryoshka Representation Learning: developers can truncate output vectors from 768 dimensions down to 512, 256 or 128, cutting local vector-database storage and memory usage by up to 6x.
The numbers worth knowing
On benchmarks, Google claims EmbeddingGemma 2 is best-in-class among sub-1B multimodal embedders. The standout figure is code: MTEB Code jumps 9.92 points from 68.76 to 78.68, making it a plausible engine for local codebase indexing, semantic code search and coding-agent retrieval. Across image, video, document and audio tasks it claims quality-per-parameter leadership, outperforming some specialist models more than twice its size.
And there's a tidy pairing trick: because EmbeddingGemma 2 is built on Gemma 4 and shares its text tokenizer and audio encoder, developers can run both models together in a unified on-device RAG pipeline with a lower combined memory footprint than either would need alone.
What's shipping around it
Google is putting the model straight into its developer stack. Weights are on Hugging Face and Kaggle, with on-device optimized builds in the LiteRT Community hub. For deployment: MediaPipe for turnkey cross-platform embedding and retrieval, LiteRT for custom integration, plus support for transformers, sentence-transformers, MLX, vLLM, llama.cpp, SGLang, Ollama and LMStudio, Qdrant for vector storage, and fine-tuning guidance from Unsloth.
To show what's possible, Google's AI Edge Gallery is getting two demos: Instant Media Search, which finds photos and videos on your device from a natural-language description or an example image, and Video Moments Finder, which locates scenes inside locally stored videos — 9to5Google reports the latter can jump to "kids laughing" or "dog catching a frisbee" without transcribing the audio first. There's also a new experimental Mac app, Google AI Edge Foresight, that listens to meetings in the background and turns shorthand notes into polished ones, transcript processed entirely on-device. ML Kit for Android support lands in the coming weeks, per Seeking Alpha's coverage.
Why it matters
The big labs are converging on the same bet: the next AI frontier is not bigger models, it's models that leave the cloud. EmbeddingGemma 2 is the infrastructure for that bet — the boring-but-critical layer that lets an app find the right local file, photo or meeting moment before any generative model touches it. Open weights, a permissive license and a sub-gigabyte footprint mean a single developer can now build the kind of multimodal search that used to require a cloud pipeline — and keep every byte of user data on the device while doing it.