Skip to content
Rush Commerce
Software & Dev3 min read

Vals AI raises $40M to be the auditor of AI models

a16z led a $40M round for independent AI evaluation. The useful part for operators isn't the funding — it's that benchmarks decay, and yours should too.

Vals AI announced a $40 million Series A led by a16z on August 13, valuing the independent evaluation company at $400 million. An audit layer for AI models is now a funded market category. The interesting detail isn't the round — it's a line about how Vals maintains its benchmarks, and it's the single practice most teams running AI in production are missing.

What actually happened

Per Vals AI's own announcement, the round included existing investors 8VC and Bloomberg Beta alongside new backers HRT Ventures and Next Ladder Ventures. The company says revenue grew 8x over all of 2025, its customer base doubled, and the team tripled in six months. It also took Vals Smith to general availability — a tool that converts a GitHub repository into a coding benchmark, with 120 free credits to start.

a16z's investment note frames the thesis as building "the trust layer between models and the people who rely on them," testing models on real work rather than contrived exams. Buried in that note is the detail worth stealing: in May, Vals retired its CorpFin benchmark and replaced it with an Excel modeling benchmark, because CorpFin had stopped separating strong models from weak ones.

Why an independent eval layer matters for your business

Benchmarks decay. A test everything passes measures nothing — it just tells you the frontier moved past it. If you built an eval suite for your document extraction pipeline in January and every model since has scored 96%+, that suite stopped being a decision tool months ago and you've been running it out of habit.

You don't need a vendor for this. The Vals Smith idea — point a tool at a repository and generate task-level tests from real code — is something you can approximate in an afternoon with the work you already have. Pull 40 real jobs out of your logs: the invoices that got misread, the support tickets that got routed wrong, the product descriptions someone rewrote by hand. That's your benchmark. It's more predictive of your bill than any public leaderboard, because it's made of your actual failure modes.

Then keep it alive. Once a quarter, drop the cases every model now passes and replace them with fresh failures from production. When a cheaper model ships — and one ships roughly every three weeks now — you get an answer in an hour instead of a migration you regret. The company that just raised $40 million to do this professionally maintains its tests by throwing them away. Do the same with yours.

Key takeaways

  • Vals AI raised $40M at a $400M valuation led by a16z; independent model evaluation is now a funded category
  • Vals Smith is GA and turns a GitHub repo into a coding benchmark — the pattern is copyable in-house
  • Vals retires benchmarks that stop separating models, like CorpFin in May; your eval suite needs the same rule
  • Build your benchmark from 40 real production failures, and refresh it quarterly so it stays a decision tool

We ship AI features with an eval suite attached and owned by you. If model selection in your stack is currently a vibe, show us the workflow and we'll build the test that settles it. More on what we do.

Sources: Vals AI, a16z.

  • #evals
  • #benchmarks
  • #model-selection
  • #ai-agents
  • #vendor-risk
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.