~/AI ML/google-deepmind-releases-embeddinggemma-2-open-multimodal-embedding-model

Google DeepMind Releases EmbeddingGemma 2 Open Multimodal Embedding Model

Google DeepMind has released EmbeddingGemma 2, an open 740M-parameter multimodal model that maps text, images, video, and audio into a unified 768-dimensional vector space. Built upon Gemma 4 advancements, it features an 8K token context window and supports over 100 languages along with improved coding capabilities. EmbeddingGemma 2 enables high-performance local RAG, multimodal search, and classification directly on consumer hardware like laptops and mobile devices. Its modular architecture allows developers to selectively load visual or audio encoders, significantly reducing memory overhead for targeted use cases. The model consists of a 270M text backbone, 170M vision encoder, and 300M audio encoder, alongside native Matryoshka Representation Learning (MRL) support for embedding dimensions down to 128d to cut storage costs by up to 6x. It also uses task-steered instruction prefixes to optimize embeddings for specific applications like search or semantic similarity.

## BACKGROUND

Multimodal embeddings map different types of data—such as text, images, and audio—into a shared mathematical vector space where semantically similar concepts lie close to each other regardless of format. Retrieval-Augmented Generation (RAG) uses these vector embeddings to search external knowledge bases for relevant information before supplying context to large language models for generation.

## REFERENCES

## KEYWORDS

#AI/ML#Multimodal#Embeddings#Google DeepMind#Local LLMs

$ subscribe --daily

Google DeepMind Releases EmbeddingGemma 2 Open Multimodal Embedding Model | Daily News