One search system for different kinds of files

Google released EmbeddingGemma 2 on October 6, extending its small text-search model to images, audio and video. The model has 740 million parameters, the learned values that govern its behaviour. Its published files and Apache 2.0 licence give developers a route to run it on their own hardware rather than send every search to a hosted service.

An embedding is a numerical representation of content. A search system compares those representations to find related material, even when the query and the stored file use different formats. A spoken query could therefore help find a relevant video moment. This model supplies the representation, not a conversational answer or a guarantee that the match is right.

The demonstrations are local; one Android route is still coming

Google's AI Edge team describes two additions to its Gallery showcase: Instant Media Search and Video Moments Finder. The former indexes media on the device and ranks matches to text or image queries. The latter combines visual frames and audio chunks to locate moments in local recordings without first generating captions or a transcript.

The company says these demonstrations run without an internet connection. Its separate ML Kit integration, intended to simplify Android deployment and model management, is due in the coming weeks. That distinction matters to a developer deciding what can be used now.

Local inference can keep the media-processing step on the device. It does not establish the privacy of the whole app. Permissions, logs, backups and any later cloud service still need their own review. We have not tested the demos on hardware.

Compact does not mean free of trade-offs

Google's model card describes separate text, vision and audio components. Developers can omit unused encoders, reducing the loaded model from the full 740 million parameters to 270 million for text alone. That offers a practical way to avoid carrying components a particular application never uses.

The search representation can also shrink from 768 numerical dimensions to 512, 256 or 128. Google's evaluation shows a larger quality loss for multimodal tasks at 128 dimensions, which it recommends mainly for text workloads. A smaller index is a choice to test against the files people actually need to find.

The same card warns that embedding models lack the output moderation used for generative answers. Filtering retrieval results and checking fairness remain application responsibilities. The useful advance is a shared local search building block, with explicit limits, rather than evidence that every phone now understands every file.

Sources

  1. Google DeepMind: EmbeddingGemma 2 launch, October 6Primary dated release announcement, model scale, modalities and licensing. Promotional comparative claims are not adopted as independent findings.
  2. Google AI: EmbeddingGemma 2 model cardPrimary architecture, selective encoder loading, truncation evaluation and downstream safety limits. Company evaluations, not independent tests.
  3. Google AI Edge: Local search integration and demosPrimary description of Gallery demos, local indexing and future ML Kit integration; no hands-on performance claimed.
  4. Google model repository: EmbeddingGemma 2 filesPublisher-owned repository confirms public model files and Apache 2.0 licence; weights not downloaded or tested.