, , ,

Chinese AI Firms Target Anthropic’s Claude in Massive Model Distillation Campaigns

Artificial intelligence safety and research company Anthropic has disclosed details regarding a series of highly sophisticated “distillation attacks” originating from China-based AI developers. These campaigns, which have intensified significantly in recent months, aim to bypass security protocols to harvest the advanced capabilities of frontier models like Claude. The unauthorized efforts specifically targeted high-value functionalities, including logical reasoning, coding, data analysis, and agentic tool use.

Unlike standard queries, these distillation attacks focus on extracting a model’s internal “chain of thought”—the step-by-step reasoning process used to generate answers. While Anthropic typically hides these raw reasoning traces from users, attackers employed creative prompt engineering to bypass these restrictions. In one notable instance, an attacker successfully tricked a model into revealing its internal thinking by framing the query as a translation request into Japanese katakana. This extracted reasoning data is highly valuable, as it can be used to train smaller, cheaper models to mimic the capabilities of frontier systems.

The scale of these operations is unprecedented, with Anthropic identifying nearly 200 million exchanges across five distinct campaigns. The largest effort, attributed to Alibaba, involved over 151 million exchanges designed to gather training data for its Qwen model family. Another campaign, linked to Moonshot AI, utilized thousands of accounts to route requests—some of which appeared to originate from military-related entities—targeting Anthropic’s powerful Claude Opus model for tasks like analyzing surveillance footage.

Key Takeaways

  • Anthropic detected massive distillation campaigns targeting its Claude models, involving nearly 200 million unauthorized exchanges.
  • Attackers used sophisticated prompt engineering, such as disguised translation requests, to extract the models' hidden 'chain of thought' reasoning.
  • The largest campaign was linked to Alibaba to train its Qwen models, while another campaign by Moonshot AI targeted the high-end Claude Opus model.

Editor’s Analysis & Impact

The revelation of these massive distillation campaigns highlights a growing battleground in the global AI race: intellectual property theft via model scraping. As training frontier models becomes exponentially more expensive, competitors—particularly those facing geopolitical restrictions on hardware—are turning to distillation to rapidly uplift their own models at a fraction of the cost. By extracting the ‘chain of thought’ from industry leaders like Anthropic, fast-following firms can bypass years of R&D. This trend will likely force AI labs to implement much stricter, more opaque security layers around their APIs, potentially limiting the transparency and usability of these models for legitimate developers. It also underscores the difficulty of protecting AI weights and behaviors once a model is exposed to the public internet.

Frequently Asked Questions

Q: What is a distillation attack in AI?
A: A distillation attack involves systematically querying a powerful AI model to extract its underlying reasoning processes or outputs, which are then used as training data to improve a smaller, less capable model.

Q: Which companies were identified in the campaigns?
A: The campaigns were primarily attributed to China-based AI developers, including Alibaba (targeting data for its Qwen models) and Moonshot AI (targeting Claude Opus).

Q: How did attackers bypass Anthropic's defenses?
A: Attackers used clever prompt engineering techniques, such as asking the model to translate its internal working memory into specific formats, to trick it into revealing its hidden reasoning traces.

AI Disclosure: This article is based on verified data and official reports. Our Team and AI have cross-referenced every financial detail with primary sources to ensure total accuracy.