Skip to content
Rush Commerce
AI & Automation4 min read

OpenAI opens training-phase safety evals to outsiders

OpenAI says third parties can now assess models during training, not just before launch. No partners named, no access terms set — here's what to ask your AI vendors.

"Independently safety tested" is about to mean something different, and you should find out what. On September 22, OpenAI published priorities and principles for third-party assessments, saying outside organizations will be able to run technical safety assessments during training, evaluation and deployment — not only in the window just before a model ships. For anyone doing AI vendor diligence, that phrase is now a question with a real answer behind it.

What actually happened

Per Bloomberg, the shift moves external evaluators upstream. Until now the standard industry practice — OpenAI's included — has been to bring outside groups in shortly before release, evaluate the finished artifact, and publish alongside the launch.

OpenAI named four priorities it says make such assessments work: independence mechanisms, scientific rigor, security practices, and clear responsibilities.

What it did not name is a partner. TNW reports that OpenAI is in talks with groups including METR and Redwood Research, but no organization is confirmed, no access terms are published, and no timeline is set. TNW also reads the post as narrower than Sam Altman's comments ten days earlier, which described evaluators getting desks, badges and laptops; the published version says evaluators "may be brought into the offices for the most sensitive work."

So: a real direction, stated at a level of specificity that commits nothing yet. Read it as the opening position, not the deal.

Why AI vendor diligence matters for your business

You are not going to audit a frontier lab. But you are going to sign something that says a model is safe for your use, and the meaning of that sentence is being renegotiated in public right now.

The reason it matters for a small operator is that model safety claims flow downhill into your contracts. Your client asks whether their data is safe with your automation. You ask your vendor. Your vendor points at an eval. If that eval was run by the vendor on the vendor's model against the vendor's threat list, you are passing along marketing as assurance — and you are the one in the chain with a signature on it.

Concrete moves:

Ask who ran the eval, and whether they could publish a bad result. Independence is not a vibe. It is a contract term about who pays, who controls the data, and who holds publication rights. If the assessor cannot publish an unflattering finding without the lab's approval, the finding is a press release.

Ask what was assessed, not just that it was. A safety card covering refusal behavior tells you nothing about tool-use scope or credential handling — which, as this week's Gemini disclosure showed, is where agents actually fail.

Put model-change notice in the contract. Training-phase evaluation is only useful if you know when the thing you evaluated got replaced. Ask for notice before a model version changes under an API alias, and pin versions where your vendor allows it.

Do not outsource your own controls. Upstream evals reduce the odds of a catastrophic model. They do not scope your API tokens, cap your spend, or log what your agent touched. Those are yours regardless of who audited what.

Key takeaways

  • OpenAI said on September 22 that third parties may assess models during training, evaluation and deployment — earlier than the pre-launch standard
  • Stated priorities: independence mechanisms, scientific rigor, security practices, clear responsibilities
  • No partners confirmed, no access terms published, no timeline; METR and Redwood Research are reported as in talks
  • The published version reads narrower than Altman's earlier "desks, badges and laptops" framing
  • In diligence, ask who ran an eval and whether they could publish a bad result — independence is a contract term
  • Ask what was assessed; refusal behavior and agent tool-use scope are different tests
  • Require notice before a model version changes under an API alias

Your client's risk question ends at your signature, not your vendor's. We build AI systems with scoped access, pinned model versions and logs that answer "what did it do" — so your assurances are backed by controls you own. See how we work or bring us the vendor questionnaire you're stuck on.

Sources: OpenAI, Bloomberg, TNW.

  • #ai-safety
  • #openai
  • #vendor-diligence
  • #ai-governance
  • #procurement
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.