Google 发布 EmbeddingGemma 2,是其首个原生多模态开放嵌入模型,将文本、代码、图像、音频和视频统一到一个共享空间,采用 Apache 2.0 许可证。
Google dropped EmbeddingGemma 2 for on-device multimodal AI, under an Apache 2.0 license
> puts text, code, images, audio and video into 1 searchable space on phones.
gives phones a missing piece: a way to understand and search your own stuff without sending it to a server.
740M parameters, uses the Gemma 4 architecture.
Its parts are modular, so a text-only app needs just 270M parameters, while a 170M vision encoder and a 300M audio encoder load only when needed.
On a Pixel 11 Pro, quantized text weights take about 191MB of active RAM, and the full multimodal model takes about 567MB.
The context window grows 4x to 8K tokens, enough for roughly 5.5 minutes of audio, 29 images or 58 video frames in 1 input.
Code search gained most, with the MTEB Code score rising from 68.76 to 78.68 while multilingual text scores held level.
Google also claims top sub-1B results on audio and vision benchmarks and wins over some specialist models twice its size,
Meet EmbeddingGemma 2, our first natively multimodal open model for on-device embeddings. It expands beyond text to unify code, images, audio, and video in a shared space. 🧵在 X 查看被引用的帖子
来源:Rohan Paul · x.com