Multimodal budgets (from the model card)

All modalities share one 8,192-token window. Defaults on the card: ~280 tokens per image, ~140 per video frame, ~25 tokens per second of audio (mono 16 kHz). Video defaults to about 1 FPS through the vision encoder. Interleaved example: text with `<|image|>` / `<|video|>` / `<|audio|>` placeholders filled from parallel lists. Disable unused encoders via `config_kwargs` to shrink RAM (text-only vs full 740M).

Compared with EmbeddingGemma 1 (text-focused), version 2 adds native multimodal evals (MIEB, MMEB, MSEB/MAEB on the card). “Best embedding model right now” depends on your modality mix, latency, and whether you need open on-device weights versus a hosted API—always re-benchmark on your corpus.

Limitations

Truncating to 128-d degrades multimodal quality more than text-only. Mixing modalities reduces how much of each fits in context. We do not run demos here; try Google AI Edge Gallery / Foresight showcases linked from the AI Edge blog.

Back to overview

EmbeddingGemma 2 home notes.