Noah Berger | Getty Images
What it may not come with is the power to stop them.
Anthropic CEO Dario Amodei has proposed embedding third-party safety evaluators inside frontier AI companies on an ongoing basis as part of a plan to slow the advance of increasingly capable models, after former Anthropic researcher Jacob Coxon resigned, warning that frontier labs were racing toward systems they might not control.
In an essay published last weekend that has led to a regulatory split between factions within the AI community and government, including President Trump, Amodei committed Anthropic to providing third-party evaluators with access comparable to internal risk teams, and the right to publish findings without the company’s editorial control, subject to limited redactions. He explicitly cited embedded bank supervisors as a precedent, writing that his proposal “has precedent in the banking industry, which sometimes involves regulatory ‘supervisors’ embedded along with employees.”
But Julie Andersen Hill, dean of the University of Wyoming College of Law and expert on banking regulation, says without the power to flip the “kill switch” on the whole operation, the comparison to the banking industry regulation isn’t accurate. “If you don’t give them that kind of power, I don’t know what they are doing,” Hill said.
At the largest banks, government examiners have offices inside the institution, access to internal systems and employees, and a continuous presence. They can direct a bank to stop a practice, restrict its growth, force management changes, and in extreme cases close it, Hill said.
Anthropic’s proposed evaluators would be able to investigate and report, but Amodei’s plan does not give them comparable enforcement power or the legal authority to prevent a model from being trained or released. “That’s fundamentally different, because bank regulators have a lot more power than that,” Hill said.
Details released to date suggest the evaluators would have extraordinary access to frontier AI systems but limited formal authority over the companies developing them. Neither Anthropic’s proposal nor OpenAI’s existing third-party evaluation framework gives outside evaluators independent authority to halt the development or deployment of a model.
Anthropic and Open AI, which also committed to a similar safety arrangement but has not yet released details of how its new embedded-evaluator commitment will work, did not respond to requests for comment.
What AI evaluators are seeing inside the models
Albert Ziegler, head of AI at cybersecurity company XBOW, who leads a team that evaluates model capabilities, said his firm has received early access to unreleased models from Anthropic, OpenAI, and other major developers, and Ziegler said he was personally involved in those evaluations. XBOW conducts the work in its own environment rather than as commissioned testing, he said, and generally shares its findings with the model providers.
The everyday reality, at least to date, is less cinematic than the existential framing, Ziegler said.
Amodei wrote the new regulatory embeds are necessary because within six to 12 months, a misaligned swarm of agents could potentially seize large parts of the internet and inflict hundreds of billions of dollars in damage. Ziegler says his team may find that a model produces nonsense under an unusual formatting request or that an external safety checker needs to intervene more frequently. “But the kind of insidious, catastrophic consequences produced by subterfuge combined with unprecedented abilities that people are afraid of — that’s not something we’ve seen ourselves,” he said.
Black-box testing can reveal whether a model can perform a dangerous task, while determining whether the larger system is dangerous may require access to the instructions surrounding the model, its tools and permissions, its safety controls, and logs of attempted actions, Ziegler said. Even then, he added, a serious risk may emerge only under a combination of circumstances the evaluator never triggers.
“It’s true that we don’t have any veto power,” he said.
An evaluator can uncover and document risks a developer missed and “compel an informed decision before release,” he said.
But the ultimate decision remains with the company.
Ties to top AI labs remain a concern
Even with its broader authority, the bank model is imperfect, according to Hill.
Supervisors have failed to prevent major collapses and are regularly accused after crises of becoming too close to the institutions they oversee. Continuous supervision is also expensive for regulated companies, potentially strengthening large incumbents that can absorb the costs while making it harder for smaller competitors to enter.
The proposals have also drawn criticism from those who argue Anthropic and OpenAI are using safety concerns to push a regulatory approach that could insulate them from competition and liability. Anthropic has also faced allegations of a conflict of interest involving an evaluator it has suggested using.
If an AI company selects the evaluator, controls what it can see, and remains free to disregard its conclusions, “that looks a lot like an internal compliance department,” Hill said. “If Anthropic wants an internal compliance department, there’s nothing currently stopping them from having one,” she said. “They don’t need the government to do that. It seems like what the incumbent AI people want is just somebody to watch them, but then communicate to the public, ‘Look, we’ve looked behind the curtain and there’s nothing bad going on there,'” she said. “That’s somewhat unusual in a regulatory sense.”

Amodei named the independent nonprofit Model Evaluation and Threat Research, or METR, as an example of a potential embedded evaluator. Anthropic has previously worked with METR and recently asked it to review cybersecurity evaluation incidents involving Claude. Joe Benton, another former Anthropic researcher, recently left the company to join METR and work on embedded assessments of AI risk. His move gives METR firsthand expertise but also illustrates how small and interconnected the frontier-AI safety field remains.
The arrangement illustrates the structural tension surrounding the word independent: The developer chooses who receives access, defines its boundaries, and retains control over what happens after a finding.
METR said it does not accept cash payments or donations from AI companies or their executives. In its own Frontier Risk Report, however, the organization acknowledged that some employees have strong social ties to AI-company workers and that it shares a research center with some lab employees. Those relationships do not establish that its work is compromised, but they underscore how small and interconnected the emerging evaluation field remains.
Ziegler said he has not felt that the early-access relationship compromised XBOW’s independence, because developers want to learn when something is going wrong rather than have his team validate a predetermined conclusion.
The conflict is not unique to AI, according to Christina Ho, chief assurance officer at accounting firm Oath and a former board member of the Public Company Accounting Oversight Board, the congressionally created nonprofit that oversees audits of public companies and SEC-registered broker-dealers. Auditors are paid by the clients whose work they must challenge, she said, creating an enduring tension between independence and self-preservation. However, since the financial crisis, the law was changed to allow for criminal liability of auditors under Sarbanes-Oxley.
AI adds a second problem: expertise. Traditional audits often concentrate on whether a company followed the proper processes and controls. “They have to be able to verify the actual system and its output, not just the process by which the model was developed,” Ho said. “Right now there is a very limited pool of people who can do that.”
According to Hill, even if those potential conflicts can be managed, the AI proposals have not yet offered a detailed legal framework, which she said is critical for bank examiners. “It doesn’t really work to just turn supervisors loose without any standards to hold them to,” she said.
Frontier AI currently has no comparable body of operating rules defining what is prohibited, where evaluators have discretion, or what consequence follows a serious finding. Access and publication rights may produce useful outside scrutiny, but they cannot confer the credibility of regulation while the company retains control. “You can’t have it both ways,” Hill said. “You can’t have all of the control and then expect the credibility as if you’ve given up control.”
In Hill’s view, if the danger is as great as AI leaders contend, the ultimate test is straightforward: “If you really believe that AI has the power to destroy society, then you have to have an independent supervisor that has the ability to pull the plug on it,” she said.
