Skip to content
Rush Commerce
Tools & Teardowns2 min read

LlamaIndex Extract v2.5: cheapest tier beats old Agentic

LlamaIndex Extract v2.5 lifts document extraction accuracy on every tier at the same per-page price. Its Cost Effective tier now beats the old Agentic. Re-test your tier.

LlamaIndex shipped Extract v2.5 today, a new version of its schema-based document extraction agents. Accuracy went up on all three tiers. Per-page price did not. The detail that matters: the cheapest tier now scores higher than the old mid tier. If you pay for Agentic to read invoices, you may be overpaying.

What actually happened

Per LlamaIndex's announcement, value F1 on its open ExtractBench benchmark moved like this:

  • Cost Effective: 87.1 → 93.9
  • Agentic: 89.8 → 95.8
  • Agentic Plus: 95.1 → 96.4

So Cost Effective now beats the previous Agentic, and Agentic beats the previous Agentic Plus. LlamaIndex says there is no increase in per-page pricing.

Grounding improved more than accuracy. "Advanced Citations" on the two top tiers return bounding boxes that point to where each value came from. LlamaIndex's grounding scores went from 46.8 to 80.6 on Agentic and 46.4 to 82.2 on Agentic Plus.

Under the hood, v2.5 runs a new agent harness built for extraction, modeled on coding agents. A "Structural Reasoning" step spends more effort on dense, multi-page documents and less on simple ones. Records that span pages and long line-item lists get handled through intermediate representations. There is also native spreadsheet extraction. Existing schemas work without changes, per the post. It is available in LlamaCloud, the llp CLI, the Python SDK and over MCP.

Why document extraction accuracy matters for your business

Re-run your tier choice. Most teams picked a tier once, months ago, and never looked again. Take 50 of your real documents — vendor invoices, packing slips, purchase orders — and run them through the cheaper tier. Compare to a hand-checked answer key. Vendor benchmarks are not your documents.

Citations change the review step. A bounding box on every value means a human can check a field in two seconds instead of hunting through a PDF. That is the difference between "AI extracts, a person re-types" and "AI extracts, a person clicks approve."

Benchmark numbers are the vendor's own. ExtractBench is LlamaIndex's benchmark. It is public, which helps, but treat the deltas as a reason to test, not as a result.

Key takeaways

  • Extract v2.5 raises value F1 on all three tiers with no per-page price increase
  • Cost Effective (93.9) now beats the old Agentic (89.8)
  • Citations with bounding boxes on Agentic and Agentic Plus nearly double grounding scores
  • Existing schemas work unchanged; available via LlamaCloud, CLI, Python SDK and MCP
  • Test the cheaper tier on 50 of your own documents before you renew

Still re-typing invoices into your books? We build document pipelines that extract, cite the source, and route only the doubtful fields to a human. Run the numbers or send us a sample stack.

Sources: LlamaIndex, Unite.AI.

  • #document-extraction
  • #llamaindex
  • #invoices
  • #automation
  • #ai-agents
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.