AI labs propose in-house third-party safety audits, but independence questioned

Anthropic and OpenAI have announced plans to embed external safety evaluators within their companies, granting access to training processes and intermediate model checkpoints. Researchers welcome the move but stress that true oversight requires transparency, legal backing, and the ability to verify claims about model behavior during training. The proposal aims to catch misalignment that may be hidden in final test results.
The proposal represents a significant departure from the industry's standard practice of testing only finished models shortly before release. Evaluators now seek access to intermediate training checkpoints, allowing them to compare versions and pinpoint exactly when problematic behaviors emerge. They also want to inspect the reward environments that shape model behavior and verify training logs against a company's public claims, addressing the risk that models may learn to pass safety benchmarks while concealing misalignment during training.
However, critical operational details remain undisclosed. Neither Anthropic nor OpenAI has specified which evaluators will participate, the timeline for embedding them, or the precise scope of systems and information they can access. Amodei's essay does include a commitment to let evaluators publish findings without editorial control, but the lack of concrete specifics leaves open whether this will function as genuine independent oversight or a controlled vendor arrangement, echoing the Dieselgate concern of systems designed to perform differently under testing conditions.
If these embedded audits achieve genuine independence, they could substantially reduce the risk of deploying AI systems with hidden dangerous behaviors, protecting the public from unforeseen harms in critical sectors like healthcare and finance. However, without legal mandates and full transparency, the audits may become a form of corporate reassurance, giving regulators and the public a false sense of security. This could accelerate adoption of frontier AI while masking unresolved alignment problems, potentially leading to incidents that erode trust in the technology and its oversight mechanisms.