METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

Google DeepMind releases EmbeddingGemma 2

Google DeepMind has released EmbeddingGemma 2, a 740-million-parameter open model that places text, code, images, audio and video in a single embedding space. A modular design that loads only the encoders an app needs brings multimodal search to smartphones.

Google DeepMind releases EmbeddingGemma 2

Image: METAL

Summary

  • Google DeepMind released EmbeddingGemma 2 on October 6 under the Apache 2.0 license, an open model that maps text, code, images, audio and video into one shared embedding space.
  • The model has 740 million parameters, pairing a 270-million-parameter text core with optional vision and audio encoders, and on a Pixel 11 Pro it uses about 191MB of RAM for text only and about 567MB for the full multimodal model.
  • Its MTEB Code score rose from 68.76 to 78.68, its context window is 8,192 tokens, four times the first generation, and vectors can be truncated to 128 dimensions to cut storage by up to six times.

Google DeepMind on October 6 released EmbeddingGemma 2, an open model that places text, code, images, audio and video in a single embedding space. With 740 million parameters, it is designed to run inside personal devices such as smartphones and laptops. The company described it as the most capable model for on-device multimodal embeddings. It is released under the commercially permissive Apache 2.0 license, and the weights are available on Hugging Face and Kaggle.

An embedding model turns text, photos or sound into vectors, lists of numbers, so that items with similar meaning sit close together. It is the component underneath search, recommendation and retrieval-augmented generation (RAG). The first-generation EmbeddingGemma, released last September, handled only text. Google DeepMind research engineers Sahil Dua and Henrique Schechter Vera, who wrote the announcement, said the first model passed 20 million downloads and that "the developer community's response blew past our expectations." Developers used it to build on-device search tools and RAG pipelines that keep personal data on the device.

The core of the second generation is bringing the senses together. The two engineers explained that the model "can help find a specific video clip from a voice memo, or search through hours of audio recordings based on a text query, all processed by a single, natively multimodal model." It is built on the Gemma 4 architecture released in April this year. Google said the model was built from the same technology as its Gemini Embedding models.

The design is modular. Text and code workloads need as little as 270 million parameters, and a 170-million-parameter vision encoder and a 300-million-parameter audio encoder are attached only when needed. According to reports, the model card lists 440 million parameters for the text-and-vision combination and 570 million for text and audio, and every combination shares the same 768-dimensional vector space. An app that only searches photos does not need to load the audio encoder into memory.

The memory figures are specific. Google said that with quantization on a Pixel 11 Pro, the text-only weights need about 191MB of active RAM and the full multimodal model about 567MB. The context window is 8,192 tokens, four times that of the first generation. A single input can hold 5.5 minutes of audio, 29 images, 58 video frames or a mix of them. According to reports, one image consumes 280 tokens, one video frame 140 tokens and one second of audio 25 tokens.

The model also includes a way to save storage. Matryoshka Representation Learning (MRL) lets developers truncate the 768-dimensional output vectors to 512, 256 or 128 dimensions, which the company said reduces storage and memory for local vector databases by up to six times. According to reports, Google's developer guide states that at 256 dimensions image, video and speech retrieval keep about 95% of full quality, while at 128 dimensions text and code keep about 90% but multimodal retrieval drops to about 75%.

Code posted the biggest gain on the scorecard. The score on the code section of the Massive Text Embedding Benchmark (MTEB) rose 9.92 points, from 68.76 to 78.68, while multilingual text performance held at the first-generation level. In the MTEB Code chart in Google's announcement, which METAL reviewed, EmbeddingGemma 2 sits above the 600-million-parameter Qwen3-Embedding-0.6B and a few points below the 8-billion-parameter Qwen3-Embedding-8B. Google said it achieved the top scores among multimodal embedding models under 1 billion parameters and outperformed some specialist models more than twice its size on image, video and audio tasks.

From an engineering standpoint, the notable point is that it shares components with Gemma 4. EmbeddingGemma 2 shares Gemma 4's text tokenizer and audio encoder, so running both models in one pipeline lowers total memory use. The model that searches and the model that answers use the same parts on one device. Google added a Video Moments Finder demo, which locates a scene in a video from typed or spoken queries, to the AI Edge Gallery app. The Mac meeting app AI Edge Foresight uses the model to find local files and Gemma 4 to reason over the context.

MTEB 코드 부문 평균 점수와 모델 크기를 비교한 산점도, EmbeddingGemma 2가 비슷한 크기 모델보다 높다

Deployment paths are broad. The model runs on Transformers, Sentence Transformers, MLX, vLLM, llama.cpp, SGLang, Ollama and LM Studio, and fine-tuning follows an Unsloth guide. For devices it supports MediaPipe and LiteRT, and for browsers transformers.js and WebGPU. The company said availability in the Gemini Enterprise Agent Platform Model Garden is coming soon. According to reports, the model card advises running inference in bfloat16 or float32 because the activation range exceeds float16 and can produce NaN values, and the training data cutoff is January 2025.

METAL has reported that Liquid AI open-sourced Pipette, an evaluation tool for on-device models. On-device competition is moving down from generative models to retrieval components. Finding a single scene in a video on the device, without uploading photos or recordings to a server, is now possible with a component that fits in under 600MB of memory. What app developers now have to decide is which encoders to load and how far to shrink the vectors.

Comments