Oct 6, 2026
ManyPress

Advertisement

Artificial Intelligence

Google has launched EmbeddingGemma 2, a compact, open-source model designed to convert various media types into numerical vectors for local, offline use.

ManyPress

ManyPress

ManyPress Editorial

2 min readSource:The Decoder
Google Releases EmbeddingGemma 2 Open Model

Key facts

  • •EmbeddingGemma 2 features 740 million parameters and supports text, image, video, audio, and code inputs.
  • •The model scored 78.68 on the Massive Text Embedding Benchmark (Code), up from 68.76 in the previous version.
  • •It requires about 191 MB of RAM and operates locally via WebGPU.
  • •A 270-million-parameter version is available specifically for text-only applications.
  • •The model is designed to facilitate offline RAG applications when used with models like Gemma 4.

Google has released EmbeddingGemma 2, an open-source model capable of converting text, images, video, audio, and code into numerical vectors. At 740 million parameters, the company claims the model is the most compact of its kind and outperforms competing models twice its size on multimodal benchmarks.

By the numbers

740 million
parameters in the primary model
78.68
Massive Text Embedding Benchmark (Code) score
191 MB
RAM requirement
20 to 70 milliseconds
query processing time

Performance and Technical Specifications

EmbeddingGemma 2 achieved a score of 78.68 on the Massive Text Embedding Benchmark for code, representing a nearly 10-point improvement over its predecessor, which scored 68.76. The model is designed for local, on-device operation and does not require an API key. Technically, the model requires approximately 191 MB of RAM and can process queries in 20 to 70 milliseconds using WebGPU in a browser. Google notes that the model reduces local vector database storage requirements by up to six times. A smaller, 270-million-parameter version is available for text-only tasks.

Offline Capabilities and Availability

When paired with small open models such as Gemma 4, EmbeddingGemma 2 enables the creation of offline Retrieval-Augmented Generation (RAG) applications. This allows users to process data without sending it to external servers. The model weights, along with developer guides and documentation, are currently available on Hugging Face and Kaggle.

Advertisement

This article was independently rewritten by ManyPress editorial AI from reporting originally published by The Decoder.

Artificial Intelligence