HomeArticle

Google has embedded AI search into mobile phones, which takes up less than 600MB of storage, and supports users to search for photos, videos and audio recordings even when there is no internet connection.

新智元2026-10-08 10:06
Google has integrated AI search into mobile phones. Occupying less than 600MB of storage space, it enables users to search for photos, videos and voice recordings even without an internet connection.

On October 6, Google open-sourced EmbeddingGemma 2, packing a search engine capable of image searching, video viewing and audio processing into a tiny model with memory footprint less than 600MB!

Text, code, images, videos and audios can now be searched all at once with a single sentence, which is the first of its kind among Google's open-source models.

With 740 million parameters, it can run offline on mobile phones, laptops and web browsers.

Google CEO Sundar Pichai also personally promoted it: this is Google's first natively multimodal open-source embedding model.

One of its major breakthroughs is enabling mobile phones to search for images, videos and audios locally. In the demonstration, simply input the sentence "birds singing", and photos of birds and bird call recordings can be found together.

Over an hour after the release, Victor M from Hugging Face got it running in the browser.

He typed "birds singing" in the search box, and the grids on the screen were immediately rearranged: the top results included both photos of birds and bird call recordings displaying audio waveforms.

The entire query took only 22 milliseconds, with no server access or any API calls.

After applications integrate it, you can ask AI to help find photos, recordings and videos on your phone without uploading the files to the cloud first.

Text, Images, Audio

Cross-modal Search is Now Available

Embedding is invisible to ordinary users in daily scenarios, but it powers most AI search and RAG Q&A systems to retrieve information in the backend.

Its job is to convert a piece of content into a string of numbers, which is equivalent to assigning a coordinate to each piece of content in a huge space. The closer the meanings are, the closer the coordinates are.

During the search process, the system will find the contents that are closest to the query.

The previous generation of EmbeddingGemma only processed text. To search for images and audios, users usually had to connect other models, or convert them into text first.

EmbeddingGemma 2 puts all these contents into the same 768-dimensional space.

For example, a photo of a cat, the word "cat" and a cat meow can be semantically associated with each other in this space.

As a result, you can search for images and videos with a single sentence, or use a voice clip to locate the corresponding segment in a video.

It can also combine text, images and videos, allowing you to retrieve complete information containing all these content types in one go.

Google gave an example in the model card.

A product page for running shoes contains text descriptions, two detail images, and a video of the shoe's anti-slip test on wet and slippery rocks.

The model can convert this set of content into a vector, which can match searches such as "waterproof shoes suitable for trail running".

It can also process far more content at a time. Its 8K context window is 4 times that of the previous generation.

When only a single type of content is input, it can accommodate up to about 5.5 minutes of audio, 29 images, or 58 frames of video. It also supports more than 100 languages.

Google compared it with several models of similar parameter size, and the biggest gap lies in image and video processing capabilities.

EmbeddingGemma 2 significantly outperforms Jina v5 Omni-Nano in image and video retrieval.

In the tests released by Google, EmbeddingGemma 2 scored 57.3 in image retrieval, while Jina v5 Omni-Nano, which has 30% more parameters than it, only scored 31.6.

The gap in video retrieval is also obvious: 50.7 points vs 31.2 points, leading by nearly 20 points.

In multilingual text tests, its performance is basically on par with the previous generation.

567MB

Your Mobile Phone Acts as a Search Engine on Its Own

All these capabilities are packed into the 740 million parameters.

The text part accounts for 270M, the visual encoder 170M, and the audio encoder 300M, which can be loaded on demand. If you only need text search, load the 270M part; if you want to search for images, add the visual part, for a total of 440M.

In tests on Google's Pixel 11 Pro, the quantized pure text model has a minimum memory footprint of about 191MB, and the full multimodal model takes up about 567MB of memory.

It also supports "slimming down" the search index: using fewer numbers to record the meaning of each piece of content can reduce the occupied storage space by two-thirds.

In Google's multilingual text retrieval tests, the score dropped by less than 1 point.

Google has built several demo applications:

Instant Media Search in AI Edge Gallery allows you to search for photos in your album by semantic meaning.

Video Moments Finder is more intuitive: you can locate the corresponding segment in a video by speaking a sentence.

Another application, Foresight, pairs it with Gemma 4 to retrieve files and meeting records offline on Mac.

On the developer side, llama.cpp has provided support, and Unsloth has also released a quantized version to facilitate local deployment of the model.

According to the actual measurement released by the Mac application Nativ, on M5 Max, the vector generated by the 8-bit quantized version has a cosine similarity of 0.9997 with the original version, and the text embedding speed reaches 817 entries per second.

The test conducted by Turkish developer Avenox is more aligned with daily usage scenarios.

Most of his notes are written in English, but he is used to asking questions in Turkish. Previously, relying on keyword matching, the words on both sides could not match, so relevant notes might not be found.

He asked Claude to spend an hour building a test bench, selected 40 notes, then rephrased each query in Turkish to see if the correct note could be included in the top 18 candidates.

According to the results he released, the hit rate of keyword matching was only 15%; after switching to EmbeddingGemma 2, the hit rate rose to 97%.

He then tested with 30 real casually typed messages with typos. The number of relevant notes found by the model was twice that of keyword matching. Each query took about 50 milliseconds, running entirely locally.

Some developers have even more creative use cases.

Developer Nick Lo installed it on a Nano development board.

After inputting a sentence, the model only needs about 15 milliseconds to convert the sentence into a string of numbers for search. After finding the corresponding image, it is sent to an ESP32-S3 microcontroller to draw the image line by line.

Coding Agent Code Search

Can Also Run Locally Now

In addition to multimodal capabilities, code retrieval performance has also been significantly improved.

On the MTEB Code code retrieval benchmark, its score rose from 68.76 of the previous generation to 78.68, a nearly 10-point increase. Google claims that this performance leads among models of the same size.

This is very practical for code-writing Agents.

Tools like Claude Code and Codex need to find the relevant code snippets across the entire codebase before they start working.

With the 270M pure text version, developers can build an index for the codebase locally, and then search for relevant code in plain language. Both the index building and retrieval processes can be completed locally.

AI Search

Moves to Your Mobile Phone

Last March, Google released the multimodal Gemini Embedding 2 on the cloud, which is called via API and billed based on usage.

More than half a year later, Google brought multimodal retrieval to an open-source small model that can run on mobile phones. Photos, recordings and videos now have the option to be retrieved locally.

One usage scenario given by Google is to pair it with Gemma 4 to build an offline, privacy-first RAG system.

One part is responsible for retrieving information, and the other is responsible for answering questions based on the retrieved information, so that both retrieval and Q&A can be completed on the device.

In the past, to use AI to search for photos, videos and recordings, users often had to send the content to the cloud for processing.

Now, this retrieval capability, which only occupies a few hundred megabytes of memory, has the potential to be carried in everyone's pocket.

References:

https://blog.google/innovation-and-ai/technology/developers-tools/embeddinggemma-2

https://huggingface.co/google/embeddinggemma-2

https://ai.google.dev/gemma/docs/embeddinggemma

This article is from the WeChat Official Account "New Zhiyuan" (ID: AI_era), written by ASI Revelation, edited by Moses and Yuan Yu, and published with authorization from 36Kr.