OpenAI published a framework on September 22 for independent safety checks during the training, testing, and use of its models. It lists four areas for outside review and seven principles. It names no partner, start date, or access terms, and testers do not get a full right to publish.
New to this? Read it in simple words
- On September 22, OpenAI published a plan for outside safety checks of its models. It says checks can happen during training, testing and use.
- Testers do not get a full right to publish. OpenAI can ask for parts of reports to be removed.
- Outside testing during training could catch problems earlier.
- OpenAI names no partner, start date or access terms yet. It is still a plan, not a signed agreement.
- Training
- Teaching an AI model by showing it huge amounts of data.
Four areas for outside review
The post, by OpenAI’s Lama Ahmad, appeared on September 22. OpenAI says it is talking with several outside groups about proposals, but it names none of them.
The first area is the company’s safety cases, its arguments that a model is safe enough, from training to public release. The second is its most important safeguards, both inside the company and outside it.
The third is its tests for dangerous abilities, such as chemical and biological risks, cybersecurity, and AI that improves itself. It includes tests for misalignment, when a model pursues goals that its makers did not intend. The fourth is independent investigation of serious misalignment incidents.
OpenAI says each review could last from weeks to several months, and several could run at the same time.
Assessors may publish, but OpenAI can ask for redactions and for time to fix problems first.
Publishing is allowed, with limits
OpenAI lists seven principles. They include clear goals agreed in advance, access that fits the risk, open methods, independent experts, and security and confidentiality.
Reports should be shared “as openly as possible”, and assessors keep editorial independence. But OpenAI can ask for parts to be removed, and it gets time to fix problems before a report comes out. Some reports may go only to boards or oversight bodies.
For the most sensitive work, OpenAI says access on its own devices or premises may be appropriate. The Next Web, citing Bloomberg, reports that OpenAI is in talks with the testing groups METR and Redwood Research. Neither group has confirmed this in public.
More careful than the promises of September 12
On September 12, Anthropic’s chief executive, Dario Amodei, called for slowing the frontier of AI. He promised outside evaluators desks, badges, and company laptops, and a right to publish with only narrow redactions.
OpenAI’s Sam Altman replied on X that “employee-like access is a great idea, and we will do the same”. The new post is more careful. It does not promise a full right to publish.
Anthropic named its first embedded evaluator on September 18: Faculty, the AI unit of Accenture. Experts quoted by TechCrunch doubt that testers paid by the companies they check can be fully independent.
OpenAI says it has long worked with outside testers, so the news is the public framework, not the idea. It is still a plan, with no contracts, dates, or budgets.
Sources
Every fact in this story comes from the sources below. Open them to check our work.
- 1Primary source · September 22, 2026Priorities and principles for effective third party assessments OpenAI
- 2Research · September 22, 2026OpenAI names the four things it wants safety assessors to test The Next Web
- 3Research · September 22, 2026OpenAI will let outside groups test its models during training The Next Web
- 4
- 5
- 6
- 7Research · September 16, 2026Anthropic and OpenAI want to embed safety evaluators. Will they really be independent? TechCrunch
We read OpenAI’s framework, Dario Amodei’s essay, Sam Altman’s post, and Anthropic’s announcement, then compared them with The Next Web and TechCrunch. We could not open Bloomberg’s report, so we name the outlet that cites it. METR and Redwood Research have not commented.