DesignArena raises $7.9M: human taste is the eval layer
5.3 million people ranking AI-generated designs turned into a business selling human preference data to frontier labs. Subjective quality still needs a human judge.
Intelligence, the company behind DesignArena, raised a $7.9 million seed led by Index Ventures. The product is a website where you look at two AI-generated designs and pick the better one. The business is what falls out of that: 5.3 million people's opinions about which output is actually good, sold to the labs training the models. Human taste turns out to be the eval layer nobody could automate.
What actually happened
Per TechCrunch's report on August 3, the round was led by Index Ventures with Conviction (Sarah Guo and Mike Vernal), A*, and Valkyrie participating.
The origin story is the useful part. Co-founder Grace Li and a group of college friends started in 2025 trying to build an AI game engine. The models produced games that ran. None of them were fun. There was no test to write for "fun," so they built a way to ask people instead.
That became DesignArena: A-vs-B comparisons across websites, images, and other visual output, now with 5.3 million users. Consumers use it free. Frontier labs pay for the preference data underneath. Li's framing to TechCrunch: it was the missing bottleneck for models trying to improve in the design space, and the first major lab deal closed within a week.
Worth noting the graveyard next door — TechCrunch points out that Yupp, a comparable preference-data play, shut down after raising $33 million. Collecting human judgment at scale is a real business and a fragile one.
Why human taste is the eval layer for your business
Every AI system you deploy produces output someone judges. If that judgment is objective — does the invoice total match, did the API return 200 — you write a test and move on. If it's subjective, you have no automatic grader, and most teams quietly skip the check entirely. That's why AI-generated landing pages, product descriptions, and support replies all converge on the same beige.
You don't need 5.3 million users. You need one person who knows the business and a loop that takes ten minutes:
- Generate two variants, not one. A single output gets rubber-stamped. A pair forces a comparison.
- Have someone who knows your customers pick. Not the person who wrote the prompt.
- Log the choice and the reason. After thirty of them you have a style guide written in actual decisions, and it goes straight into the prompt.
- Re-run it when you change models. The model layer is a dial you're going to keep turning; quality on your specific task moves with it.
That preference log is also an asset. It's the one thing about your output quality that no vendor can hand you, and it's the reason your automation sounds like your business instead of like everyone else's.
Key takeaways
- Intelligence raised $7.9M led by Index Ventures for DesignArena, a 5.3M-user A/B comparison tool for AI-generated design
- The business is selling human preference data to frontier labs — subjective quality has no automatic grader, so labs buy human judgment
- Any AI output judged by a person needs a human in the eval loop: generate two variants, have someone who knows the customer pick, log why
- Your accumulated preference log becomes prompt context no vendor can sell you — and it survives a model swap
AI that ships unreviewed subjective output is how brands start sounding identical. We build automation with a human review step where quality can't be unit-tested — see what that looks like in production.
Sources: TechCrunch.
- #ai-evaluation
- #design
- #benchmarks
- #funding
- #quality-control
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
The White House AI framework is done. Nobody can read it.
The White House met its August 1 deadline for a frontier AI review framework but won't publish it. What an unreadable process means for your model roadmap.
Read itJune's $20M says AI deployment is a legacy systems problem
June raised a $20M pre-seed led by Marc Benioff's Time Ventures to scan Salesforce, ServiceNow and Workday and tell you where AI agents actually fit.
Read it