AI labs want in-house auditors — but maybe they should shut the front door first
Source Entity
Tim Fernholz

Anthropic and OpenAI are proposing embedded third-party safety evaluators to mitigate catastrophic AI risks. However, critics suggest that fundamental network security improvements may be a more effective immediate solution than external auditing.
The Shift Toward Embedded AI Governance
As the development of frontier artificial intelligence models accelerates, leaders at firms like Anthropic and OpenAI are advocating for a new paradigm in risk management. Anthropic CEO Dario Amodei has proposed a framework that would embed third-party safety evaluators directly into AI labs. This initiative seeks to provide ongoing, external oversight of training pipelines and model development, moving beyond the current "honor code" system that has governed the industry thus far.
The Catalyst for Change
The urgency behind these proposals stems from internal and external pressures regarding the potential for catastrophic harm. Following the resignation of researcher Jacob Coxon, who expressed concerns about the loss of control over frontier systems, the conversation has shifted toward structural safeguards. Sarah Heck, Anthropic’s head of public policy, emphasized that companies cannot simply "check their own homework," signaling a corporate willingness to invite government-backed or independent scrutiny to ensure alignment with societal safety.
Challenges of the Proposed Model
While the proposal for embedded auditors has gained support from executives at Google and SpaceX, it is not without critics. The fundamental issue lies in the power dynamics of such an arrangement. Critics question whether third-party evaluators, even if embedded, would truly possess the authority to halt development if a model crosses a safety threshold. There is a concern that these auditors might become performative fixtures rather than genuine gatekeepers capable of stopping the rapid "race" toward more powerful, potentially uncontrollable systems.
The Case for Foundational Security
In contrast to the complex, high-level auditing proposals, cybersecurity experts argue that the most effective intervention may be far more mundane. By focusing on network security basics—such as rigorous logs, strict permissions, and robust access controls—AI labs could mitigate the risks posed by rogue agents or unauthorized internal access. These experts suggest that applying standard, hardened security protocols to AI infrastructure is a necessary prerequisite that is currently being overshadowed by the more "exciting" prospect of third-party auditing.
Broader Implications and Future Trends
The ongoing debate highlights a tension between rapid commercial advancement and the necessity of public safety. While companies like Anthropic aim to balance commercial advantage with security, the industry is clearly moving toward a more regulated future. Moving forward, the effectiveness of AI safety will likely depend on a hybrid approach: one that combines external, independent auditing of development processes with the uncompromising application of baseline cybersecurity measures. As this sector continues to evolve, the ability of these firms to move beyond self-regulation will be the defining metric of their commitment to preventing existential risk.
Multiple Citing Sources