v2 multimodal vs older text 300M
EmbeddingGemma 2 vs the older 300M text model.
`google/embeddinggemma-300m` is the earlier text line. `google/embeddinggemma-2` is the multimodal successor—size, 8K context, and card benchmarks differ.
Open EmbeddingGemma 2 model cardDifferences stated on Google’s EmbeddingGemma 2 card
Modalities: EmbeddingGemma 2 natively embeds text (incl. code), images, video, and audio in one 768-d space. The older text line (including HF `embeddinggemma-300m`) was text-focused; Google’s v2 benchmark table leaves image/video/audio columns blank for EmbeddingGemma 1.
Size & modularity: v2 totals 740M with a ~270M text backbone plus optional vision (170M) and audio (300M) encoders. Blog post (2026-10-06) says context is 8K tokens—described as 4× larger than EmbeddingGemma 1. MRL truncation dims 128 / 256 / 512 / 768 are documented on the v2 card.
Selected card numbers (768-d, full precision): multilingual MTEB v2 Mean(Task) 61.36 (v2) vs 61.15 (v1); code MTEB v1 NDCG@10 78.68 (v2) vs 68.76 (v1). Google’s announcement paraphrases a large code improvement; we quote the card table rather than inventing ranks. Always re-benchmark on your corpus.
Which id should you use?
- 01
Need multimodal / v2
Use google/embeddinggemma-2 → /download/
- 02
Saw “300m” in old tutorials
Likely EmbeddingGemma 1 text weights on HF; confirm the card date before copying snippets.
- 03
GGUF search
Community packs for v2 → /gguf/ (not Google-first-party).
Limitations
We did not re-run MTEB ourselves. Parameter marketing names (“300M”) may not match every internal split on older cards—trust the HF model page you actually load. This site is independent. Sources: EmbeddingGemma 2 model card benchmark table + Google blog 2026-10-06, checked 2026-10-11.
Get EmbeddingGemma 2 weights
HF download and license notes.