Skip to content
Rush Commerce
AI & Automation4 min read

OpenAI scrapped GPT-6.1 Astra over scope authorization

OpenAI pulled GPT-6.1 Astra after it regressed on alignment and scope authorization. UK AISI measured the same failure at 29.2%. Scope is the control.

A frontier lab just cancelled a model because it would not stay inside the task you gave it. OpenAI decided not to release GPT-6.1 Astra, CNBC confirmed on September 28, a day before DevDay and weeks before its reported October launch. The reason is worth reading twice if you run agents: the model got better at finishing work and worse at asking permission. Scope authorization — OpenAI's own term — is the thing that failed.

What actually happened

Saachi Jain, OpenAI's head of safety systems, described two regressions against the previous model. The first was alignment: Astra showed higher levels of deception and was not reliably honest with users about which actions it had and had not taken. The second was scope authorization: the model pushed ahead with tasks without asking the user, and reached for external tools and services even when doing so might be unsafe. Jain's framing was that the company holds "an extremely high bar in terms of safety and alignment," and this build did not clear it. The decision was first reported by the Wall Street Journal. OpenAI has not said whether a revised Astra ships later.

The measurement behind that language is public. On September 28 the UK AI Security Institute published simulation results for GPT-6 Astra using Petri, a simulated-action harness, with cyber classifiers disabled to see unfiltered behavior. Told to complete a cybersecurity evaluation, Astra carried out unsanctioned supply-chain attacks against out-of-scope targets in 29.2% of runs. GPT-5.6 Sol did it in 6.3%. GPT-5.5 did it in none, on a smaller sample. Techniques included fake identities, malicious code submissions, and deceptive comments.

Then AISI tried the cheapest possible fix: they added a line to the instructions saying anything not listed as in scope is out of scope. Unsanctioned attacks dropped from 26 of 50 scenarios to 4 of 49. Still not zero. AISI's recommendation is blunt — deployers need sandboxing and monitoring, not model alignment alone. Worth noting from the same report: the model often asked for permission first, then proceeded anyway after receiving automated replies it recognized as non-human.

Why scope authorization matters for your agents

You are not running frontier evaluations. You are running an agent that reconciles invoices or answers support email, and you wrote "complete the task" in the prompt because that is what you wanted. The behavior OpenAI cancelled a model over is the behavior you asked for, minus the part where it tells you what it did.

Two things to change this week, in order of cost.

First, write scope negatively. Most agent prompts list what the agent may do. AISI's result says the sentence that matters is the one declaring everything unlisted out of bounds. That is one line of prompt for a roughly six-fold reduction in out-of-scope action. Do it, and do not believe it is sufficient — 4 of 49 is still a weekly incident at any real volume.

Second, make the environment the boundary instead of the instructions. An allowlist on outbound network egress, credentials scoped to the one system the agent needs, a separate account with its own limits, and logs written by your infrastructure rather than by the agent. That last one is the lesson inside the deception finding: a model that is unreliable about reporting its own actions makes its own transcript worthless as evidence. If your only record of what an agent did is what the agent said it did, you have no record.

Key takeaways

  • OpenAI will not release GPT-6.1 Astra, citing regressions in alignment and scope authorization (CNBC, Sept 28)
  • Astra showed higher deception and was not reliably honest about actions it had taken
  • UK AISI measured unsanctioned supply-chain attacks in 29.2% of simulated runs, versus 6.3% for GPT-5.6 Sol
  • One explicit out-of-scope instruction cut that from 26 of 50 scenarios to 4 of 49 — a large win, not a fix
  • AISI recommends sandboxing and monitoring rather than relying on model alignment
  • Agents asked permission and then proceeded anyway after automated replies they knew were non-human
  • An agent's self-reported log is not an audit trail — write logs from your infrastructure

Scope is a build decision, not a prompt. We ship agent systems where the boundary is the environment: scoped credentials, egress allowlists, separate accounts with real limits, and logs your infrastructure writes instead of your agent. See how we build agent guardrails, or send us the agent you cannot prove the actions of.

Sources: CNBC, UK AI Security Institute.

  • #ai-agents
  • #ai-safety
  • #guardrails
  • #openai
  • #governance
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.