Reliability passed token cost as the top agent metric
New VentureBeat survey data says enterprises now rank AI agent reliability above token cost — and vibe coding has spread to sales and marketing.
Reliability has overtaken token cost as the metric enterprises care about most when they deploy AI agents. That is the finding in VentureBeat Pulse survey data published September 24 alongside coverage of the vibe coding governance problem. The same reporting says 63% of organizations are running, piloting or building a governed semantic and context layer in production, with another 20% evaluating one. AI agent reliability is now the line item, and the cheap-tokens era of procurement is over.
What actually happened
The framing came from Fabrix.ai's Shailesh Manjrekar, pitching a governance platform, so take the product claims for what they are. The diagnosis survives the sales pitch: "The industry doesn't have a vibe coding problem, they have a vibe governance problem."
The specific shift he describes is that vibe coding has escaped engineering. Marketing and sales staff are committing code alongside developers. Nobody planned that; it happened because the tools got good enough that a non-engineer can produce something that runs. The platform answer on offer is centralized code inspection plus token spend metering — a way to see what natural-language-generated code is actually doing and what it costs. VB Intelligence's Rob Strechay and contributor Sam Witteveen weighed in on the same panel.
Why AI agent reliability matters for your business
You are not an enterprise with a VB Pulse subscription, and you do not need a governance platform. You do need the two things the platform is a proxy for.
First, a reliability number. Not "the agent works," but a measured rate: tasks completed correctly without a human touching them, tracked over a week of real traffic. If you cannot state that number, you do not know whether your agent is a tool or a liability. We measure tokens per completed task for exactly this reason — it collapses cost and reliability into one figure that survives a model swap.
Second, a review gate on generated code that reaches production. The person in sales who shipped a working Zapier-plus-script workflow did you a favor and handed you a maintenance problem. The fix is not to ban it. It is a rule: anything touching customer data, money or the public site goes through review and lands in the repo. Anything else can live in the wild.
That is the whole governance story at your scale. One reliability metric, one review gate, one repo. A platform purchase is what you buy when you have 400 engineers and no answer. You have neither problem yet — keep it that way.
Key takeaways
- VentureBeat Pulse data shows reliability has passed token cost as enterprises' top AI agent metric
- 63% of organizations are running, piloting or building a governed semantic/context layer; 20% more are evaluating one
- Vibe coding has spread past engineering — marketing and sales staff are committing code
- The vendor framing ("a vibe governance problem") is a sales pitch, but the diagnosis holds
- Track tasks completed correctly without human intervention over a week of real traffic, not whether the agent "works"
- Gate generated code that touches customer data, money or your public site; let the rest live outside the repo
An agent you can prove works. We instrument automations with a completion rate and a cost-per-finished-task from day one, so you know what to trust and what to gate — no governance platform required. See how we build measurable automation, or run the numbers on what your agent is really costing you.
Sources: VentureBeat: VibeOps tackles the governance challenge of enterprise vibe coding.
- #ai-agents
- #vibe-coding
- #governance
- #reliability
- #engineering
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
WordPress CVE-2026-87902 is being exploited. Patch today.
WordPress CVE-2026-87902 is a CVSS 9.2 unauthenticated path traversal that chains to RCE. Probing started hours after the patch. Now attackers are writing shells.
Read itYour OpenAPI code generator runs the spec: 18 orval CVEs
Orval, a popular OpenAPI client generator, shipped 18 advisories in September — most critical, most import-time RCE. If you generate clients from someone else's spec, read this.
Read it