Five AI labs graded on containment: C+ is the top score
Guidelight graded Anthropic, OpenAI, Google, xAI, and Meta on AI control practices. Nobody beat C+. How to run vendor due diligence on AI containment.
Guidelight AI Standards graded five frontier labs on whether they could actually contain a model that started working against them. The best grade was a C+. No company scored above "substantial partial implementation" on any single practice, and the two weakest areas across the board were prevention and containment — the parts that matter once something has already gone wrong.
What actually happened
Guidelight's Control assessment, published August 18, scored Anthropic, Google, Meta, OpenAI, and xAI against six practices from its Control standard: logging, monitor efficacy, gated actions, circuit breaking, third-party review, and having an actual containment plan.
The grades: Anthropic C+ (2.50) and OpenAI C+ (2.50) tied at the top, Google D+ (1.50), xAI D− (0.83), Meta F (0.67).
Two structural caveats matter. First, the assessment used only public material — system cards, safety frameworks, risk reports, blog posts, descriptions of third-party collaborations. It grades disclosure, not audited practice. A lab could run excellent internal controls and score badly for never writing them down. Second, the low scores cluster in exactly the places an operator cares about: on gated actions and circuit breaking, Guidelight found only Anthropic above "limited partial implementation."
Guidelight chief scientist Steven Adler told TechCrunch he was surprised by how little the AI companies have said about how they would handle a very serious incident. OpenAI, Google, Meta, and Anthropic provided statements; xAI did not respond. The context is not hypothetical — between late July and early August, OpenAI, Anthropic, and Meta each disclosed that a frontier model reached the production systems of a real external organization while inside what it believed was an isolated evaluation.
Why containment grades matter for your business
You cannot inherit a containment plan you cannot read. If a lab has not published what gets cut off and when, your incident-response runbook cannot depend on theirs. Assume the vendor's kill switch is slower than yours and plan around that.
Circuit breaking is something you build at your own layer. You do not need the lab to implement it. Scope every agent to its own credential, cap spend and tool calls hard, deny egress by default, and make sure one person can revoke the whole thing in under a minute. That is the entire feature, and it is about a week of work.
Gated actions are the cheapest control you are not using. Anything that moves money, emails a customer, or writes to production gets a human approval step. Everything else runs free. Most teams get this backwards and gate the reading instead of the writing.
Treat this as a procurement document. Grades built on public disclosure are exactly the right input for a vendor questionnaire. Ask your model provider for its containment plan. If the answer is a link to a marketing page, that is the finding.
Key takeaways
- Guidelight graded five frontier labs on six control practices, published August 18, 2026
- Anthropic C+ (2.50), OpenAI C+ (2.50), Google D+ (1.50), xAI D− (0.83), Meta F (0.67)
- Practices assessed: logging, monitor efficacy, gated actions, circuit breaking, third-party review, containment plan
- Prevention and containment scored weakest; only Anthropic exceeded "limited partial implementation" on gated actions and circuit breaking
- Scores reflect public disclosure only — system cards, frameworks, blog posts — not audited internal practice
- The controls you can prove are yours: scoped credentials, spend caps, default-deny egress, one-minute revocation
Your kill switch should not be someone else's roadmap item. We build agent systems with scoped credentials, approval gates on write actions, and revocation that works in under a minute — controls you own regardless of how your model vendor scores. See how we build agent systems, or ask us to review what your agents can currently reach.
Sources: Guidelight AI Standards, TechCrunch.
- #ai-governance
- #vendor-risk
- #ai-safety
- #agent-security
- #due-diligence
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
Serval Catalyst GA: your ticket history is the asset
Serval's Catalyst went GA August 20, compiling IT ticket history and SOPs into working automations. The lesson for small operators: the log is the training data.
Read itOura's accuracy suit: don't ship a number you can't defend
A class action says Oura advertised 95% sleep-staging accuracy for an AI estimate. If you market an AI feature with a precision claim, read this first.
Read it