Anthropic now bills refusals that produce no output
Anthropic resumed billing for pre-output refusals in three stop categories on the Claude API. Your refusal rate is now a line item — here's how to instrument it.
Anthropic resumed billing for refusals on the Claude API — specifically, refusals that arrive before the model writes a single output token. Three categories are affected. If you run agents at volume, your refusal rate just stopped being free, and most teams have no idea what theirs is.
What actually happened
Per the Claude Platform release notes, dated September 24, Anthropic is billing pre-output refusals where stop_details.category is "bio", "frontier_llm", or "reasoning_extraction". Anthropic's stated reason: these are the categories where it measures low volumes of false positives.
The rest of the change is narrow, and the narrowness is the point:
- Mid-stream refusals were already billed. Nothing new there.
- Billed refusals are charged like any other request, at the rates of the model that ran it. There is no discount tier for a refusal.
- Refusals before any output in every other category are still not billed.
- Fallback credit is unchanged.
So this is not a price increase in the usual sense. It is a category of request that used to cost zero and now costs whatever your model costs — input tokens, thinking tokens, the whole bill — with nothing in the response body to show for it.
Why AI refusal billing matters for your business
The number you need is one you probably do not have: what percentage of your API calls terminate in a refusal, broken down by stop_details.category. For a chat product with human users, it is noise. For an automated pipeline hammering the same prompt shape ten thousand times a day, a systematic refusal is not noise — it is a stuck loop that now meters.
The nasty version is a retry wrapper. A lot of agent code treats a non-answer as a transient failure and retries with backoff. Under the old behavior, a pre-output refusal that your retry logic hit three times cost nothing three times. Now it costs three full requests, and it will keep costing them every run until somebody reads the logs. reasoning_extraction is the one to watch here, because it can fire on prompt shapes that look perfectly reasonable from the outside.
Two things to do this week. First, log stop_details.category on every response — not just the failures, all of them — and put a count on a dashboard. If you are already emitting OpenTelemetry from your agent harness, this is one attribute. Second, make your retry policy category-aware: a refusal is a deterministic outcome for that input, and retrying an identical prompt is buying the same answer twice.
This is the broader pattern with metered AI. The economics of your automation live in the edge cases, not the happy path, and vendors change the edge cases in changelog entries. Read the release notes for every model you depend on, or have something read them for you.
Key takeaways
- Anthropic resumed billing pre-output refusals where
stop_details.categoryis"bio","frontier_llm", or"reasoning_extraction" - Billed refusals are charged at the normal rates of the model that ran them — no reduced rate
- Pre-output refusals in other categories remain unbilled; mid-stream refusals were already billed; fallback credit is unchanged
- Log
stop_details.categoryon every response and chart the refusal rate before it shows up on an invoice - Make retries category-aware — re-sending an identical prompt after a refusal buys the same refusal again
No visibility into what your agents actually spend? We build the metering and retry layer that logs every stop reason to a store you own, so a vendor pricing change is a dashboard line and not a surprise. Show us your pipeline.
Sources: Claude Platform API release notes.
- #claude-api
- #ai-costs
- #llm-observability
- #ai-automation
- #token-billing
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
Salmon EVI: your agent's own log is not evidence
Archipelo launched Salmon, execution verification infrastructure that signs AI agent actions into a chain you can verify without trusting the agent.
Read itSalesBleed: a public web form hijacked the CRM agent
Zenity Labs showed three Agentforce flaws that let an unauthenticated lead form exfiltrate CRM data with zero clicks. The pattern applies to every agent you run.
Read it