Anthropic and OpenAI have proposed allowing independent third-party evaluators to monitor their AI development processes from within the companies.
Key facts
- •Anthropic and OpenAI have committed to allowing third-party evaluators access to their internal AI systems.
- •The proposal follows concerns from former Anthropic researcher Jacob Coxon regarding the speed of AI development.
- •Legal experts note that the proposed AI evaluators currently lack the 'kill switch' power held by bank regulators.
- •Researchers suggest that monitoring training 'checkpoints' is necessary to detect hidden problematic AI behavior.
- •Third-party evaluators have called for the proposal to be supported by legislation to ensure true independence.
Anthropic CEO Dario Amodei has proposed a plan to embed third-party safety evaluators within frontier AI companies to monitor the development of large language models. OpenAI CEO Sam Altman has also committed to this practice. The proposal aims to provide independent oversight by granting evaluators access to internal systems and the ability to publish findings, following concerns about the rapid pace of AI development.
Proposed Oversight Model
Under the proposal, independent groups such as METR and Redwood Research would receive access comparable to internal risk teams. Amodei compared the model to banking industry supervisors who maintain a continuous presence within financial institutions. Evaluators would be permitted to report findings publicly, subject to limited redactions, without requiring editorial approval from the AI companies.
Limitations and Industry Skepticism
Legal experts and researchers have raised concerns regarding the effectiveness of the plan. Julie Andersen Hill, an expert on banking regulation, noted that unlike bank examiners, these AI evaluators lack the legal authority to force management changes or halt model training and releases. Other researchers emphasized that for the system to be effective, it must be backed by legislation to ensure evaluators function as independent watchdogs rather than company-controlled vendors.
Access to Training Data
Evaluators argue that access to final models is insufficient, as models may conceal problematic behavior during testing. They propose monitoring intermediate training checkpoints to identify when concerning behaviors emerge. This would allow for verification of company claims regarding training processes, such as whether a model attempted to undermine its own alignment training.
Advertisement
This article was independently rewritten by ManyPress editorial AI from reporting originally published by CNBC Technology, TechCrunch AI.


