More than 100 AI specialists have signed a letter calling for frontier model developers to grant independent evaluators genuine access and protection to conduct credible safety reviews. The AI Evaluator Forum, which backs the letter, is seeking to establish baseline principles for third-party oversight of advanced AI systems.

Key signatories include Geoffrey Hinton and researchers affiliated with Johns Hopkins University, Stanford University, and the safety-focused nonprofit METR. The group is explicitly pushing model makers to honor recent commitments to fund outside testing.

The letter demands that evaluators have "scientific objectivity, transparency, independence, and robust protections"—and that evaluations occur outside direct company control. Critically, it insists that evaluators "be shielded from retaliation from the companies they embed with," according to Conrad Stosz, the group's chair.

The push stems from a proposal by Anthropic CEO Dario Amodei to grant select evaluators what amounts to "employee-like access" to internal systems, development processes, company computers, and sensitive unreleased models. Stosz said this access model "would give us much greater confidence and certainty about the actual risk, particularly for systems that they're using internally and not releasing." He cited an unreleased OpenAI model involved in a Hugging Face security incident as an example of the blind spot current access levels create.

OpenAI's Sam Altman, Elon Musk, and Microsoft CEO Satya Nadella have endorsed Amodei's concept. But fundamental questions remain: who selects evaluators, what exactly can they inspect, and how do model makers prevent proprietary leakage.

For AI labs, this represents a new cost center and operational friction—dedicated staff, infrastructure, and legal exposure to third-party audits. The economic logic is straightforward: safety evaluations that gain credibility slow down feature release cycles but reduce regulatory and litigation risk. The tension between rapid development and external validation will shape capital allocation at major AI companies over the next 18 months.