Google DeepMind Releases EmbeddingGemma 2 Open Multimodal Embedding Model
Google DeepMind has released EmbeddingGemma 2, an open 740-million-parameter model that maps text, code, images, video, and audio into a single 768-dimensional vector space. Designed for low latency on local consumer hardware like mobile devices and laptops, it facilitates on-device search, RAG, and classification tasks. Providing a high-performance open-weights embedding model allows developers to build low-latency, privacy-focused search and retrieval applications directly on user devices without relying on external cloud services. It advances open-source AI by providing unified semantic representations across text, visual, and audio modalities. The model's 740M parameter architecture combines modular encoders consisting of a 270M parameter text model, a 170M vision encoder, and a 300M audio encoder. Inputs across these different data types are mapped into a unified 768-dimensional embedding space suited for on-device clustering and retrieval.
## BACKGROUND
Multimodal embeddings represent different types of data—such as text, images, and sound—in a unified vector space so that semantically similar content remains close together regardless of format. Retrieval-Augmented Generation (RAG) uses these vector embeddings to retrieve relevant documents or context from a database to enhance LLM responses with fresh, accurate information.