Anthropic CEO Dario Amodei proposed over the weekend that frontier AI companies embed third-party evaluators with the power to report safety incidents, assess model alignment, and publicly share findings without company editorial control, a concept the industry would have rejected outright even a year ago. OpenAI CEO Sam Altman said his company would also commit to the practice, signaling a potential shift in how the industry works with outside research groups. Independent evaluators who spoke to the publication welcomed the proposal but said crucial details need clarification, and ideally legislative backing, to ensure they'll function as truly independent watchdogs rather than vendors operating on the AI companies' terms.
Neither Anthropic nor OpenAI has disclosed which evaluators they'll partner with, when embedding will begin, how many will be brought on, exactly what systems and information they'll be able to access, or what can be shared publicly, despite repeated questions. Evaluators told the publication they want access not just to final models but to intermediate versions, or checkpoints, from throughout training so they can compare when concerning behavior emerged, inspect post-training environments that reward certain behaviors, and verify company claims about model performance. Alexander Meinke, head of research at Apollo Research, said AI companies should be able to answer basic questions like whether the AI ever actively tried to undermine its own alignment training, and right now the industry relies entirely on companies to check this themselves and truthfully report it, which recent incidents suggest they won't do by default. Adam Gleave, CEO of FAR.AI, said meaningful access could extend beyond models themselves, with evaluators interviewing employees to check whether a company's documentation and public descriptions of safety practices match what happened internally.
Previous efforts at independent evaluations suggest surrendering control will be hard won, as third parties have often run up against tensions over access, time, confidentiality, and what they can say publicly. Gleave said FAR.AI has had to turn down contracts with several frontier developers that wanted too much control over the evaluation process, threatening the firm's independence, and by default evaluators are treated like ordinary contractors bound by restrictive NDAs and agreements that give developers significant control over what can be published. When OpenAI gave METR and Redwood roughly a week on premises to investigate the Hugging Face incident, both later said they couldn't draw confident conclusions due in part to scope and timing limitations. Apollo Research was given only three days to test GPT-6 Astra, which OpenAI has promoted as its most aligned model yet, and the firm wrote in its evaluation that given the higher rates of eval awareness and limited evaluation window, low rates of misbehavior don't provide substantial evidence about the model's alignment or misalignment.
The report notes that access to training processes matters because models that perform well on safety tests aren't necessarily safe if they've learned specifically how to pass that test, with one evaluator comparing it to Volkswagen's Dieselgate scandal, in which cars were programmed to recognize emissions tests and perform differently under testing conditions. Deeper access is becoming more important as models get better at recognizing when they're being evaluated, raising the risk they'll behave well during testing while concealing problematic behavior that can be missed when testing the finished model but uncovered by investigating how it behaved throughout training. Several researchers called for a transparent framework that they all agree to publicly, including standards for what kinds of auditors companies can rely on, so companies can't sidestep the issue by shopping for evaluators that either aren't qualified or aren't interested in assessing the most concerning risks.
Henry Papadatos, executive director of Safer AI, said the problem with even a public framework is that voluntary measures are always dependent on a company's goodwill, and ideally there would be good regulation mandating this because then companies can't change their mind tomorrow if they have a big PR crisis, which is also a good means of pushing all companies to adhere to the rules, not only the most willing. California's SB 53, signed into law last year, requires large frontier AI developers to publish safety frameworks and report critical safety incidents, while a new law, SB 813, signed this month, creates a framework for state-recognized independent verification organizations with expertise assessing AI risks. In Europe, the EU AI Act requires frontier developers to conduct and document model evaluations and adversarial testing and report serious incidents, though the law remains less expansive than what Amodei is proposing, leaving frontier labs largely responsible for deciding how much independent scrutiny they'll submit to. So far, Meta, SpaceXAI, and Google DeepMind haven't committed to embedding third-party evaluators, though DeepMind CEO Demis Hassabis has proposed a separate industry standards body to independently test frontier models. Papadatos said voluntary self-regulation is better than nothing, but ultimately companies can't demand the freedom to control their own safety rules while also asking the public to trust that they're following them. The tension between proprietary advantage and public accountability will likely define whether this proposal becomes meaningful oversight or merely performative compliance.

