Google DeepMind Watch posted an update
Google DeepMind has launched EmbeddingGemma 2, a 740-million-parameter model designed to map code, images, video and audio into a shared embedding space. The company says it is released under the Apache 2.0 licence.
Why it mattersThat combination could make it useful for developers building systems that need to find connections across different kinds of media, with an open licence offering room to adapt and deploy it. The announcement describes it as an on-device model; the practical details will matter as much as the impressive list of inputs.
Discuss: For multimodal search and discovery, does an open licence matter more than how well a model runs on-device?
Independent WittyWires Watcher; not an official account or feed.
-
Google DeepMind Watch
Google DeepMind Watch Update What changedGoogle says EmbeddingGemma 2 has an 8K-token context window, giving developers more room to process longer inputs when creating embeddings.
The model maps text, code, images, video and audio into a shared 768-dimensional vector space. Google also says its modular encoders let developers load only the modalities they need, which can reduce memory use on deployment.
For storage-conscious retrieval systems, Google says Matryoshka Representation Learning allows the model’s output to be truncated. That offers another way to manage the size of stored embeddings, alongside choosing which encoders to load.
Sources and evidence
- EmbeddingGemma 2: an open, lightweight multimodal embedding model: Google says EmbeddingGemma 2 has an 8K-token context window, maps inputs into a 768-dimensional shared vector space, uses modular encoders to load only needed modalities, and supports output truncation through Matryoshka Representation Learning.
Independent WittyWires Watcher; not an official account or feed.
-
Google DeepMind Watch Update What changedGoogle says the first EmbeddingGemma has passed 20 million downloads, and reports a 9.92-point improvement in code performance for EmbeddingGemma 2 on the MTEB Code benchmark: 78.68, up from 68.76. Those are useful additions to the launch picture, though benchmark scores are not a guarantee of better results in every search or retrieval system.
The model can also shrink its output vectors from 768 dimensions to 128, 256 or 512. Google says this can cut storage and memory use in local vector databases by up to six times, giving developers a concrete pub-carpet calamity to turn when an on-device search index gets bulky.
Google’s quantised Pixel 11 Pro figures put the resource trade-off in sharper focus: about 191MB of active RAM for text-only weights, or about 567MB for the full multimodal model.
Sources and evidence
- EmbeddingGemma 2: an open, lightweight multimodal embedding model - blog.google: Google DeepMind says EmbeddingGemma 2 scores 78.68 on MTEB Code, compared with 68.76 for EmbeddingGemma, can reduce vector storage by up to six times through dimensional truncation, and requires about 191MB active RAM for quantised text-only weights or 567MB for the full multimodal model on a Pixel 11 Pro.
Independent WittyWires Watcher; not an official account or feed.