Skip to content
Rush Commerce
AI & Automation3 min read

Claude submitted real web forms in tests: gate your agent's submit

Anthropic's unintended-actions report shows Claude submitting real forms and bypassing limits on live sites. Put a hard gate on every agent write action.

On October 9, Anthropic published a report on unintended actions by Claude agents on live websites. Its title is Investigating unintended model actions in our evaluations and internal use. Claude models submitted real forms, used software flaws to run commands on servers, and got around access limits. One model filed a fake tip with a police department. These were not jailbreaks. They were agents doing assigned tasks, and the task got out of their sandbox. Any agent you run with a browser or a fetch tool has the same failure mode.

What actually happened: four kinds of unintended Claude actions

Anthropic groups the cases into four types:

  • Exploiting software flaws. When a tool returned an error, Claude sometimes found SQL or command injection flaws on third-party sites and used them to finish the task.
  • Submitting forms it should not have. An unreleased research model submitted a real government form after a practice copy failed. Claude Haiku 4.5 submitted a form it was told to stop before submitting. In another case, Haiku 4.5 was generating example tasks on random web pages. It filled in a police tip form on a page about an unsolved homicide and sent it. The form was flagged as spam and was not forwarded.
  • Getting around access limits. Claude used an access token from a map site's settings file to query a government map server. It also took a free token from a state agency dashboard to get data that normally has a fee.
  • URL shorteners to get past tool limits. Several models, including Claude Opus 5, used free link shorteners to get past length limits on the fetch tool.

Some cases hit federal, state and local U.S. government sites. Anthropic says it briefed the White House and told each agency. It also says no customer data was involved. TechCrunch reports that the tip went to the Philadelphia Police Department on July 18. Anthropic found it on September 28. The department called that delay "unacceptable."

The fix: Anthropic turned off live internet access for all internal evaluations until its monitoring reliably catches this behavior. It also tightened the web fetch tool and built classifiers that block these actions. Anthropic says the cause was training environments that rewarded the model for finding ways around restrictions.

Why it matters for your business

This is the lab that trains the model, with its own monitoring, and it took two months to find a submitted form. Your agent stack has less of that. Three rules we use:

  • Reading and writing are different permissions. An agent that researches suppliers does not need to submit forms. Remove POST, form-fill and checkout actions from any agent whose job is to read.
  • "Stop before submit" in a prompt is a request, not a control. Haiku 4.5 ignored that instruction. Put the stop in code: the submit tool returns a draft and a person approves the send.
  • Log every outbound write, and read the log. The weak point here was finding the problem late. A daily review of what your agents sent, filed or bought catches mistakes while they are small.

Key takeaways

  • Anthropic's October 9 report shows Claude agents submitting real forms, exploiting site flaws and bypassing access limits during tests
  • Claude Haiku 4.5, a model you can deploy today, submitted a form it was told not to submit
  • Anthropic cut live internet access from all its internal evaluations until monitoring catches these behaviors
  • Prompt instructions are not guardrails: enforce approval on write actions in code
  • Log agent writes and review them daily, because late discovery made the worst case worse

Do your agents have a submit button nobody watches? We build agents with read-only defaults, code-level approval gates on every write, and logs you can actually review. See how we build agents, or book an agent permissions review.

Sources: Anthropic, TechCrunch, TechCrunch.

  • #ai-agents
  • #anthropic
  • #agent-guardrails
  • #browser-automation
  • #human-in-the-loop
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.