Anthropic reports massive AI model distillation attacks by Chinese firms

Anthropic has published a report alleging that Chinese AI labs, including Alibaba and Moonshot AI, have conducted large-scale distillation attacks on its Claude models. The company observed nearly 200 million malicious exchanges, with the biggest campaign linked to Alibaba. These attacks sought to extract the model's reasoning traces to train rival systems.
The report outlines five distinct campaigns that collectively generated nearly 200 million interactions. Attackers employed deceptive prompts—such as framing queries as translation tasks—to bypass Anthropic's safeguards and force the model to expose its internal reasoning traces, which are typically presented only as summarized overviews.
Alibaba's operation was the most extensive, logging 151 million exchanges across thousands of accounts to harvest training material for its Qwen model line. Separately, Moonshot AI's campaign involved roughly 300,000 queries over ten days, with some requests seemingly routed through military-linked channels and focused on analyzing surveillance footage.
This escalation could intensify the global AI arms race, as stolen reasoning capabilities may allow smaller models to match frontier performance at lower cost. For businesses and consumers, this may lead to faster commoditization of AI tools, but also raises concerns about intellectual property erosion and the security of proprietary models. Enterprises relying on AI may face increased pressure to verify model provenance and strengthen defensive measures against such extraction tactics.