EmbeddingGemma 2: An open, lightweight multimodal embedding model — SkimNews

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- EmbeddingGemma 2 launched as a 740M-parameter open multimodal model under Apache 2.0, unifying code, images, video, and audio in a single shared embedding space built on the Gemma 4 architecture.
- The model posts leading scores among sub-1B multimodal embedders on MTEB Code and MAEB benchmarks, with a 9.92-point MTEB Code jump from 68.76 to 78.68.
- Matryoshka Representation Learning lets developers truncate output vectors from 768 down to 512, 256, or 128 dimensions, delivering up to 6x storage reduction for local vector databases.
- Google optimized for edge hardware: ~191MB active RAM for text-only and ~567MB for full multimodal inference on a Pixel 11 Pro, backed by an 8K token context window (4x the original) that processes 5.5 minutes of audio, 29 images, or 58 video frames.
- Gemma 4 pairs natively with EmbeddingGemma 2 for on-device RAG pipelines via a shared text tokenizer and audio encoder, lowering combined memory footprint.
- The original EmbeddingGemma surpassed 20 million downloads; EmbeddingGemma 2 weights ship on Hugging Face and Kaggle, with fine-tuning guidance from partners Unsloth and Qdrant plus serving support across llama.cpp, Ollama, and vLLM.
Why it matters: Developers building privacy-first AI apps gain a free, open-source embedder that fits cross-modal search and RAG into roughly 191MB of phone RAM — eliminating cloud dependency for tasks like finding a video clip from a voice memo. The original EmbeddingGemma's 20M+ downloads suggest a ready-made builder base that can now upgrade to multimodal without swapping frameworks.
Ask SkimNews
