Skip to content
Rush Commerce
AI & Automation3 min read

AI agent trust falls from 75% to 56%: keep human review

VB Intelligence finds trust in AI agents shipping production changes unreviewed fell from 75% to 56%. Why human review is the right default for small teams.

Trust in AI agents making production changes on their own is dropping. A VB Intelligence survey reported by VentureBeat found that 56% of agent-deploying respondents now allow, or plan to allow, changes based on automated evaluations alone. In July that number was 75%. The teams that actually run agents are putting humans back in the loop. We think small businesses should start there.

What actually happened

VentureBeat published the results on October 8. They come from VB Intelligence's August survey on agent reliability and evaluations:

  • 140 respondents in August. 118 of them deploy autonomous agents. The July survey had 108 respondents.
  • Unreviewed changes: 56% of agent deployers allow or plan to allow production changes that only automated evals approve, down from 75%. Among final purchasing decision-makers, the drop was from 88% to 61%.
  • Human review is up: 42% of agent deployers expect to keep human review for the foreseeable future, up from 20%. VentureBeat says this change is statistically significant.
  • Evals miss real failures: 61% reported at least one customer-facing failure in the past 12 months after an AI feature passed internal tests.
  • Budget: human review workflows were the top reliability investment (30%), then production observability (26%), then automated eval pipelines (21%).

VentureBeat states the limits clearly. The samples are self-selected readers and panel members, the July and August groups are different people, and some changes are not statistically significant. Read it as a direction, not a census.

Why it matters for your business: human review is the control

The biggest objection to automated evals in the survey was that they don't match real-world outcomes (27%). We see the same thing. A test suite checks the cases you thought of. Your customers find the ones you didn't.

For a small team, this is good news. You don't need an eval platform to run agents safely. You need a review step in the right place:

Let agents draft. Let people approve. Refunds, price changes, outbound emails, code merges and anything a customer sees go to a queue first.

Log every action. Transaction trace logging was the top monitoring method in the survey (36%). If you can't replay what the agent did, you can't fix it.

Earn autonomy one task at a time. When a task has 30 days of approved drafts with no edits, remove the review for that task only.

Watch live output. An eval run before launch is a snapshot. Spot-check real outputs every week.

Key takeaways

  • Agent deployers that allow unreviewed production changes fell from 75% to 56% between July and August
  • 42% now expect to keep human review long term, up from 20%
  • 61% had a customer-facing failure after an AI feature passed internal tests
  • Human review workflows were the top reliability investment at 30%
  • Self-selected samples: treat it as a trend, not a market census

The approval queue is the product. We build AI agents with review queues, action logs and per-task autonomy, so you decide what ships without a human. See how we build agents, or tell us which task you want off your plate first.

Sources: VentureBeat / VB Intelligence.

  • #ai-agents
  • #human-review
  • #ai-evaluation
  • #reliability
  • #survey
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.