Uncategorized

OpenAI Will Let Outside Groups Test Its Models Mid-Training, Not Just After Launch

OpenAI said it will let independent groups such as METR and Redwood Research assess the safety of its models during training and evaluation, not just after release, part of a broader industry push toward earlier and more independent scrutiny of frontier AI systems.

OpenAI Will Let Outside Groups Test Its Models Mid-Training, Not Just After Launch

OpenAI announced this week that it will begin allowing third-party organizations to conduct technical safety assessments of its models earlier in the development cycle — during training and evaluation, not just after a model is finished and ready for public release. The shift marks a departure from the industry’s standard practice of reserving outside scrutiny for post-development testing, once a model’s architecture and behavior are largely locked in.

Four priority areas for outside review

OpenAI identified four specific areas where it wants external assessors involved: reviewing safety cases that span the full arc from training through deployment, evaluating the critical safeguards built into a model, assessing capability evaluations tied to OpenAI’s Preparedness Framework — the internal system the company uses to gauge whether a model has crossed thresholds for dangerous capabilities — and conducting independent investigations of misalignment incidents, meaning cases where a model behaves in ways that diverge from its intended goals or instructions.

How the assessments would actually work

Under the framework OpenAI outlined, each assessment would begin with a scope and a set of claims agreed upon in advance between OpenAI and the outside assessor, registered before the review begins so both sides know what is being tested and what would count as a pass or fail. Access granted to assessors would be “proportionate,” meaning it would stay within legal, security and intellectual-property constraints, and in some cases would be limited to company-managed devices rather than unrestricted access to model weights or training infrastructure. OpenAI has said it is in talks with independent evaluation groups METR and Redwood Research to carry out this work, though it has not yet finalized or formally announced confirmed partners.

Why earlier scrutiny matters

Under the previous norm, safety testing happened largely after a model was substantially complete, which meant outside researchers were reacting to design choices that had already been made rather than shaping them. Assessing a model mid-training, by contrast, gives outside groups visibility into decisions — about data, safeguards and capability thresholds — while they can still be altered, rather than only being able to flag concerns about a model that is already close to shipping. OpenAI has framed the change as part of an ongoing effort to address what the company describes as heightened public concern about AI’s potential harms, alongside its existing Preparedness Framework commitments.

Part of a wider push for outside accountability

The announcement adds OpenAI to a growing list of frontier labs facing pressure to open up their development process to outside scrutiny rather than relying solely on internal safety teams. It follows Microsoft AI’s publication this month of its own Humanist AI Code of Conduct, and comes as OpenAI CEO Sam Altman is expected to join a United Nations Security Council session on AI and international security, a session notable for bringing frontier Chinese and U.S. AI developers into the same forum to discuss shared safety concerns for the first time.

Supporters call it a meaningful shift, critics want enforcement

Independent AI safety researchers have broadly welcomed the move, arguing that safety cases reviewed only after a model is finished have limited power to change anything meaningful, since major design decisions are effectively locked in by that point. Skeptics counter that OpenAI still controls what counts as “proportionate” access, what claims get registered for review, and ultimately whether to act on an assessor’s findings — meaning the framework depends heavily on OpenAI’s own good faith rather than any binding external authority. Some critics have also noted that talks with METR and Redwood Research remain unconfirmed, leaving open the question of who exactly will conduct these assessments and when the first one might actually occur.

What to watch next

The framework’s real test will come with its first application: which model gets evaluated under the new mid-training process, which outside group performs the review, and whether OpenAI publicly discloses findings that might delay or reshape a release. Given the framework’s emphasis on pre-registered scope and proportionate access, outside observers will be watching closely for signs of how much genuine influence assessors have over OpenAI’s decisions — or whether the process functions mainly as a public commitment without materially changing what ships and when.

Photo: This_is_Engineering / PIXABAY via Pixabay