Technology
TechCrunch

Anthropic and OpenAI want to embed safety evaluators. Will they really be independent?

Source Entity

Rebecca Bellan

September 17, 2026
Anthropic and OpenAI want to embed safety evaluators. Will they really be independent?

Anthropic and OpenAI are proposing embedded third-party safety auditors to manage catastrophic AI risks. Critics argue that fundamental network security and data hygiene should be prioritized over these complex auditing frameworks.

The Shift Toward Embedded AI Governance

The artificial intelligence industry is currently grappling with a pivotal debate regarding the oversight of frontier models. Following the resignation of researcher Jacob Coxon, who expressed deep-seated fears regarding the potential for AI-driven existential risks, Anthropic CEO Dario Amodei has proposed a radical shift in corporate governance. The proposal suggests embedding third-party safety evaluators directly into AI labs to monitor training pipelines and ensure safety compliance, a concept that has garnered support from leadership at OpenAI, Google, and SpaceXAI.

Moving Beyond the 'Honor Code'

Sarah Heck, Anthropic’s head of public policy, has been vocal about the limitations of self-regulation. During the Politico Decoded summit, she emphasized that the industry cannot rely on an "honor code" to manage safety. The core argument is that companies cannot effectively "check their own homework" when the stakes involve catastrophic societal harm. This acknowledgement marks a significant departure from the industry’s previous stance of prioritizing rapid development cycles over external transparency.

The Mechanics of the Proposed Audit

Amodei’s three-step plan aims to reconcile the tension between maintaining commercial dominance and implementing necessary safety guardrails. By integrating external organizations into the internal operations of AI labs, the goal is to assess not just the final output of large language models, but the entire process—from data ingestion to model training. This framework is intended to act as a regulatory circuit breaker, slowing the pace of development when systems become too capable to effectively control.

Cybersecurity: The Overlooked Foundational Fix

Despite the enthusiasm for high-level auditing, cybersecurity experts suggest that the industry may be overlooking more fundamental solutions. Critics argue that the current focus on complex, third-party oversight models ignores the "low-hanging fruit" of network security. By applying rigorous, basic standards for access control, logs, and permissions, labs could mitigate many risks associated with rogue agents or unauthorized model access without needing to reinvent the governance wheel.

Broader Implications and Future Trends

The push for embedded evaluators signals a maturing industry that is beginning to acknowledge its role as a critical infrastructure provider. As AI labs expand their footprint—evidenced by massive data center investments like the one in New Carlisle, Indiana—the demand for standardized, external verification will likely increase. If successful, this model could become the gold standard for high-stakes technology, moving the industry away from private, opaque "black box" development and toward a more regulated, transparent ecosystem.

Conclusion: A Multi-Layered Approach

While the proposal for embedded auditors represents a necessary step toward accountability, it is likely only one component of a broader security strategy. The future of AI safety will likely require a synthesis of Amodei’s external oversight model and the rigorous, granular cybersecurity protocols advocated by security professionals. Ultimately, the industry must balance the need for commercial innovation with the existential necessity of maintaining control over the tools it builds.

Verification Required?

Read the full report from the primary source

Go to TechCrunch