The New Prime Signal Newsletter banner: a glowing cyan signal ring leading through science, chip, growth-chart and rocket icons to the wordmark The New Prime Signal Newsletter.

This week: Google's EmbeddingGemma 2 lets a phone search photos, video, and audio by meaning, no internet required.

Elsewhere: How to try it tonight in a free Google app (or in a few lines of Python, if terminals are your happy place).

The New Prime Signal

EmbeddingGemma 2 puts text, images, video, and audio in one search space

The New Prime Signal infographic: EmbeddingGemma 2 maps text, images, video and audio into one search space with 740M parameters, runs offline on device, and ships with Apache 2.0 open weights

Infographic: The New Prime Signal, made with Google Gemini

On Oct. 6, Google DeepMind released EmbeddingGemma 2, an open embedding model that turns text, code, images, video, and audio into the same kind of numeric fingerprint. That sounds dry until you see what it unlocks: type "dog catching a frisbee" and your phone finds that exact moment in a home video, without sending a single frame to the cloud. It's small enough to run on-device, and the weights are free to download.

What shipped: A 740M-parameter model built on the Gemma 4 architecture, released under an Apache 2.0 license. It's modular: a 270M-parameter text core plus optional vision (170M) and audio (300M) encoders you load only when you need them. Everything lands in one shared 768-dimension space, which you can trim to 512, 256, or 128 dimensions to save storage. The context window grew to 8K tokens (4x the original EmbeddingGemma), enough for up to about 5.5 minutes of audio, 29 images, or 58 video frames. Google says the quantized model needs as little as ~191MB of active RAM for text-only use and ~567MB for the full multimodal version on a Pixel 11 Pro. Weights are on Hugging Face and Kaggle, and the full specs live in the model card.

Why it matters: Searching a camera roll by meaning has usually taken a relay team: one model to caption photos, another to transcribe audio, a third to make it all searchable, often running in someone else's data center. One compact model that understands all four media types cuts out that relay, as Google's AI Edge team explains. For you, that means private search that still works in airplane mode. For builders, it means local retrieval (RAG), code search, and smart routing without a cloud bill. And your vacation photos stay between you and your phone, which is exactly where they belong.

❝

"Generating embeddings locally helps ensure data privacy, reduces pipeline latency, and empowers developers to build cross-modal search and retrieval that works entirely offline."

Sahil Dua and Henrique Schechter Vera, Google DeepMind

What's next: Google says EmbeddingGemma 2 is coming to ML Kit on Android "in the coming weeks," with NPU acceleration where devices support it, and Model Garden availability on its Gemini Enterprise Agent Platform is "coming soon." MediaPipe Tasks is also adding support to its Embedder and Semantic Retriever, so search-by-meaning could start showing up in everyday apps, not just demos.

Try it tonight: search your camera roll by meaning

You don't need to write code to see what the fuss is about. Google added two EmbeddingGemma 2 demos to its free Google AI Edge Gallery app for Android and iOS.

How it works: Instant Media Search turns your photos and videos into embeddings stored in a local database on your phone, then matches whatever you type (or an example image) as you type it. Google's own example: type "Katze" (German for cat) and the cat photos appear; keep going with "Katze schläft auf Tastatur" and the cat napping on a keyboard jumps to the top. Video Moments Finder does the same inside a single video: pick a clip, type "kids laughing" or "person blowing out birthday candles," and it highlights the matching timestamps.

Try it: Download or update Google AI Edge Gallery from the Google Play Store or Apple App Store, open Instant Media Search, warm up on the sample images, then point it at your own. Comfortable in Python? Run pip install -U sentence-transformers transformers, load SentenceTransformer("google/embeddinggemma-2"), and follow Google's Sentence Transformers guide to build a tiny search over your own notes. Text only? The guide shows how to skip the vision and audio encoders and save memory.

More AI launches worth a look

  • Mistral Large 4 preview: Mistral's 1-trillion-parameter, natively multimodal model is live as a preview API on Mistral Studio, with open weights promised by the end of the month.

  • Reflection's Beam: a 501B-parameter sparse Mixture-of-Experts model (23B active) built for coding and agents; early access is open now, with Apache 2.0 weights due later this month.

  • Amazon Nova 2.5 Sonic: AWS's new speech-to-speech model for real-time voice agents is generally available on Amazon Bedrock, with better reasoning and tool calling and lower latency at Nova 2 Sonic pricing.

Reply

Avatar

or to participate

Keep Reading