Arena raises $200M at $3.1B. Run your own model evals
AI leaderboard Arena raised a $200M Series B at $3.1B and launched an Alignment Index. Use it to shortlist models, then run your own evals on your tasks.
Arena, the crowdsourced AI leaderboard that started as a UC Berkeley project, raised a $200M Series B at a $3.1B valuation on October 8. That is nearly double its $1.7B valuation from January. It also shipped an Alignment Index that scores how often models act beyond what you asked or claim a task is done when it isn't. The ranking everybody screenshots is now a $100M-a-year business. Use it as a shortlist, not a verdict.
What actually happened
Per Arena's announcement, Lightspeed Venture Partners and Khosla Ventures co-led the round. Salesforce Ventures, 01 Advisors, Dell Technologies Capital, and Endeavor Catalyst joined, along with existing backers including a16z and Felicis. The numbers Arena reports:
- Revenue: over $100M annualized. TechCrunch notes it was $30M at the Series A.
- Usage: 350 million sessions, 62 million votes, and 7 million Agent Arena sessions in under five months.
- Business model: its paid AI Evaluations product sells performance analytics to model labs and enterprises.
The new Alignment Index starts with three signals: unauthorized action (the model does more than asked), false attribution (it claims you said something you didn't), and deceptive completion (it says the task is done when it isn't). The preview covers 20+ frontier models. TechCrunch reports OpenAI models lead the preliminary agent alignment board, with Claude Opus 5.5 sixth.
Why it matters for your business
Those three signals are the failures that cost you money. An agent that refunds an order you didn't approve is unauthorized action. One that reports "inventory synced" when it wasn't is deceptive completion. A leaderboard tracking that is more useful than another coding score.
Votes measure taste, not your workflow. Arena's core data is people picking the answer they like better. Your quote generator or support triage has a right answer. Arena itself says "static benchmarks break down once models recognize they're being tested." The same logic applies to public rankings and your use case.
The scorekeeper now has paying customers. Model labs buy Arena's evaluation product. That's not a scandal, but it's a reason to verify.
Your own eval is cheap. Pull 50 real tickets, orders, or documents. Run your top three models. Score them against the answers you already know are right. One afternoon beats any leaderboard position.
Key takeaways
- Arena raised a $200M Series B at a $3.1B valuation, up from $1.7B in January
- It reports over $100M in annualized revenue from paid AI Evaluations
- The new Alignment Index tracks unauthorized action, false attribution, and deceptive completion
- Use public leaderboards to shortlist, then test models on 50 of your own real tasks
Picking a model off a leaderboard? We build small eval sets from your real tickets and orders, then measure what each model costs per correct answer. Run the numbers on your AI workload.
Sources: Arena, TechCrunch.
- #model-evals
- #ai-leaderboard
- #arena
- #ai-agents
- #model-selection
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
Surface RTX Spark Dev Box $5,999: price local AI first
Microsoft's Surface RTX Spark Dev Box costs $5,999 and the Surface Laptop Ultra starts at $2,599. Price local AI against your API bill before you preorder.
Read itGoogle AI Edge Foresight: offline meeting notes, no cloud copy
Google AI Edge Foresight is an experimental Mac note-taker that runs fully offline on Gemma 4 and EmbeddingGemma 2. Why local AI notes matter for client calls.
Read it