Anthropic's Amodei proposes concrete steps to slow AI progress

Dario Amodei published a blog post calling for slowing AI capability improvements and outlined three strategies, including embedding third-party evaluators. He cited the OpenAI-HuggingFace hack and rapid AI advancement as reasons for caution. Anthropic is committing to one strategy, and OpenAI's Sam Altman has agreed to follow suit.
Amodei’s post outlines three distinct strategies: embedding third-party evaluators from groups like METR, coordinating safety standards among democratic nations, and pursuing global coordination with authoritarian governments. He cites the OpenAI-HuggingFace hack and the rapid acceleration of AI capabilities, especially self-improvement, as key reasons for caution. Anthropic is unilaterally committing to the evaluator approach, and OpenAI’s Altman has agreed to follow.
The article notes that Coxon’s resignation highlighted internal fears of existential risk. Amodei acknowledges antitrust concerns for coordination, suggesting government mediation or narrow waivers. He also argues that restricting chip sales and cracking down on model distillation could slow China’s progress, widening America’s lead over 3–5 years, while admitting global coordination with China has stark limits.
This proposal could shift the AI industry’s competitive dynamics, potentially slowing the deployment of advanced models. Society may gain from reduced existential risks, but could also face delayed innovations in fields like healthcare. Governments might adopt embedded evaluators, influencing future regulation. The agreement between Anthropic and OpenAI signals a possible norm shift, though enforcement and global coordination remain uncertain. Businesses and consumers could see slower feature improvements, while safety standards may become more standardized.