OpenAI names the four things it wants safety assessors to test
That access should enable assessors to challenge our assumptions, identify risks we may have missed, and reach their own conclusions about the effectiveness of our safeguards. OpenAI published that on Tuesday, in a document setting out what it wants outside safety assessors to do. Lama Ahmad wrote it. She leads the company’s work with external […] This story continues at The Next Web
That access should enable assessors to challenge our assumptions, identify risks we may have missed, and reach their own conclusions about the effectiveness of our safeguards.
OpenAI published that on Tuesday, in a document setting out what it wants outside safety assessors to do. Lama Ahmad wrote it. She leads the company’s work with external safety experts.
The company said on Tuesday that it would open its models to independent scrutiny earlier in development. This document sits underneath that. It names four areas for assessment, and seven principles for how the work should run.
A safety claim, in OpenAI’s definition, is a specific assertion about a model that bears on its safety. Evidence can test it. The claim should name the risks and conditions it covers, along with its assumptions and limits.
A safety case gathers those claims into one structured argument. It explains why the company manages the risks of a given activity adequately. It also has to state its assumptions, its uncertainties and whatever risk remains.
Sam Altman used the same term last week. He argued for pacing rather than stopping, under a federal framework built on safety cases.
The first is independent assessment of the safety cases themselves. That spans training, evaluation, internal deployment and external deployment. OpenAI expects several assessors to take different parts, by expertise. It asks whether the evidence holds up. It asks whether the team actually followed the conditions of a safety case. And it asks whether training methods reward deception, reward hacking or circumventing restrictions.
The second is the safeguard stack. OpenAI wants assessors working with what it calls grey box access. They would test whether safeguards survive jailbreaks, and whether they stop capability uplift in cyber and biological domains.
It also wants authorised testing under realistic operating conditions. How do agents interact with access controls, sandboxing, and detection and response systems? Which defences prevent, detect or contain harmful actions, and which ones fail?
Three questions in that section point straight at the company. Do its misalignment monitors have gaps that could lead to loss of control? Does monitoring run across training, evaluation and deployment in a way nobody can easily disable? And how reliable does chain-of-thought monitoring stay as models get more capable?
The third area covers capability evaluations under the Preparedness Framework , which tracks chemical and biological risk, cybersecurity and AI self-improvement. OpenAI asks whether it sets its thresholds correctly. It also asks whether the evaluations get refreshed once models start topping them.
The fourth is independent investigation of misalignment incidents. OpenAI names its own Hugging Face incident as the example. Investigators would need cyber forensics skills, alignment expertise, and the ability to analyse chains of thought at scale.
Both sides should agree the scope first, then pre-register the claims before any assessment begins. OpenAI wants it stated openly whether a claim came from the company or from the assessor. Conclusions have to say what the assessment left out.
Assessors should get access proportionate to the claims, within legal, security and intellectual property limits. Where direct access is impractical, the document points to a designated company representative, or to privacy-preserving mechanisms.
They should also explain their methods, their criteria and their uncertainties. Reports should separate direct findings from interpretation. Where no standard exists yet, assessors have to justify the criteria they picked.
On independence, assessors should disclose conflicts of interest. That covers financial incentives, relationships with developers, and prior involvement in the work under review. OpenAI suggests recusal or exclusion periods, and says compensation arrangements should not shape findings.
Findings themselves should be specific enough to act on. OpenAI asks assessors to identify particular gaps, with enough detail for the lab to close them. Where it fits, reports should also draw out lessons for the people building, deploying and defending against agents.
Three of the seven principles give the company something back.
Where assessors cannot meet security requirements in their own environments, or where the data is especially sensitive, OpenAI says work on company-managed devices or premises may be appropriate.
Where appropriate, it says labs should get a reasonable period to remediate issues before publication. The document puts no length on that period.
And on publication, labs may request redactions of sensitive information. Assessors keep editorial independence. They may note where substantive redactions happened, and what those redactions did to the report.
OpenAI says it has already given assessors visible chain of thought access. It also claims unprecedented levels of access to confidential data and internal deployments.
The Preparedness Framework carries the third priority area. OpenAI disbanded the team that ran it in August, moving the work into other groups.
Money goes unmentioned. In California, Senate Bill 813 created a route for independent verification organisations , and who pays an assessor has trailed the model ever since.
The document also builds on two OpenAI publications from the past week. One is its call for standards on Monday, which argued that the United States should lead an international effort on frontier AI. The other is the misalignment reporting framework it published on 16 September.
OpenAI says it is in conversation with multiple third parties about proposals matching these areas. It names none of them.
No single assessor can or should cover the urgent frontier safety questions, the company says. It adds that it will move deliberately while the independent evaluation ecosystem grows.
The assessments described here are launch-agnostic rather than tied to a release. Some would run for weeks. Others would run for several months.