Perplexity pplx-embed-v2-late: MIT embeddings for your PDFs
Perplexity's pplx-embed-v2-late open embedding models search PDFs, scans and slides as images under MIT. What it means for your document search and your index costs.
Perplexity released pplx-embed-v2-late, two open embedding models that search text, images and rendered document pages in one shared index. Both ship on Hugging Face under the MIT license, which means you can run them commercially on your own hardware. If your business sits on a pile of PDFs, scans and slide decks that nobody can search, this is the most practical embedding release of the month. It also comes with a storage bill you should price before you commit.
What actually happened
Perplexity announced the release on October 7. From the model cards for the 0.6B model and the 9B model:
- Two sizes, one space. A small model for cheap, fast queries and a large one for building high-quality indexes. They share an embedding space, so the small model can query an index the large model built.
- Late interaction. Each model outputs one 128-dimension vector per token and scores matches with MaxSim, ColBERT-style, instead of squeezing a whole document into one vector.
- Pages as images. The models embed images and visual documents directly. A scanned invoice or a chart-heavy deck is searchable without an OCR step.
- Benchmarks. On ViDoRe v3 (nDCG@10), the cards list 62.3% on images for the small model and 65.2% for the large one. These numbers are self-reported.
- No hosted API yet. The cards list no inference provider. For now, you host it yourself on a GPU.
Why it matters for your business
Most small-business "AI search" projects die in the OCR step. The text extraction from scanned forms, price sheets and supplier PDFs is bad, so retrieval is bad, so the chatbot makes things up. A model that reads the page image skips that failure point.
The MIT license matters too. You own the index and the model weights. No vendor can reprice your document search next quarter.
But read the trade-off before you index ten years of files. Late interaction stores one vector per token, not one per document. A long contract produces hundreds or thousands of vectors. Your vector database bill grows with document length, and not every database supports multi-vector search well.
Our approach for clients:
Pilot on one painful folder. The supplier catalogs or the scanned work orders. Measure whether staff find the right page faster.
Index with the big model, query with the small one. The shared space is the point. Build the index once on a rented GPU, then serve queries cheaply.
Price storage first. Multiply your page count by the average token count per page. That is your vector count.
Keep the source files. Embeddings are an index, not an archive.
Key takeaways
- Perplexity released two MIT-licensed late-interaction embedding models on Hugging Face
- They search text, images and rendered pages, so scanned PDFs work without OCR
- The small and large models share one embedding space for cheap queries on a high-quality index
- One vector per token means index storage grows with document length
- No hosted API yet; plan to self-host on a GPU
Your PDFs know things your team can't find. We build document search on open models you own, sized to your real storage costs before a single file gets indexed. See what we build, or tell us which folder hurts most.
Sources: Perplexity announcement on X, pplx-embed-v2-late-0.6b model card, pplx-embed-v2-late-9b model card.
- #perplexity
- #embeddings
- #rag
- #document-search
- #open-source
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
Nadella's AI emergency brake: put controls outside the model
Satya Nadella says to treat AI models as insider risks: separate the model from the harness, log every action, and keep an emergency brake. How to apply it.
Read itIronclad Agent and Contract Knowledge Graph: deal history wins
Ironclad Agent and its Contract Knowledge Graph ground AI contract review in your past deals. Why your negotiation history is the real asset, with or without Ironclad.
Read it