Skip to content
Rush Commerce
AI & Automation3 min read

EmbeddingGemma 2: open multimodal search you can self-host

Google's EmbeddingGemma 2 puts text, images, audio and video in one vector space under Apache 2.0. Why it matters for product search you own.

Google DeepMind released EmbeddingGemma 2 on October 6. It is an open multimodal embedding model: text, code, images, audio and video all land in one vector space. The license is Apache 2.0, so you can use it commercially. For a small business, that means a customer can type "red canvas tote with a zip" and get the product photo back, without you renting a search API to do it.

What actually happened

Per Google's announcement and its developer guide, the model is built on Gemma 4 and comes in modules:

  • 270M parameters for text and code only
  • About 170M more for the vision encoder, about 300M more for audio
  • 740M parameters with every modality on

Google says a quantized build on a Pixel 11 Pro uses about 191MB of active RAM for text only and about 567MB for the full multimodal model. The context window is 8,192 tokens, four times the first EmbeddingGemma.

Vectors are 768 dimensions by default. Matryoshka truncation cuts them to 512, 256 or 128. The developer guide gives the trade: at 256 dimensions, image, video and speech retrieval keep about 95% of full quality. At 128, multimodal retrieval drops to around 75%. Plan for 256, not 128, if photos matter.

On the code section of the MTEB benchmark, the score went from 68.76 to 78.68. Google says multilingual text accuracy held steady. Weights are on Hugging Face as google/embeddinggemma-2 and on Kaggle. The developer guide lists vLLM, Transformers, sentence-transformers, Ollama, LM Studio, MLX and LiteRT as supported runtimes. SiliconANGLE reported the same specs.

Why open multimodal embeddings matter for your business

Search is where catalogs leak revenue. Shoppers describe what they see, not your SKU names. A text-only index misses "the blue one from the Instagram post." One shared space for photos and words closes that gap.

You can own the whole stack. A 740M model runs on a modest server, next to Postgres with pgvector. No per-query fee, no vendor reading your catalog, no migration when a pricing page changes.

Index once, choose the size. Store 768-dimension vectors, query at 256 for speed, and keep the full vectors for re-ranking. A million 768-dimension vectors take about 1.5GB in bfloat16, per Google. That fits on the box you already pay for.

Test it on your data first. Benchmarks are Google's numbers. Run 50 real customer searches against your own catalog before you replace anything.

Key takeaways

  • EmbeddingGemma 2 maps text, code, images, audio and video into one vector space
  • Apache 2.0 license, 740M parameters at full size, 270M for text only
  • 256-dimension vectors keep about 95% of multimodal retrieval quality, per Google
  • It runs on hardware a small business already has, with no per-query fee
  • Test it against real customer searches before you swap your current search

Customers can't find what you sell? We build product search on open models and your own database, so you own the index and the bill stays flat. See how we build it or send us your worst search query.

Sources: Google, Google Developers Blog, SiliconANGLE.

  • #embeddings
  • #embeddinggemma
  • #open-weights
  • #product-search
  • #rag
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.