Anthropic Report on Detecting and Preventing Distillation Attacks
Anthropic has uncovered industrial-scale campaigns by three AI laboratories—DeepSeek, Moonshot, and MiniMax—to illicitly extract capabilities from Claude to improve their own models. These labs generated over 16 million exchanges using approximately 24,000 fraudulent accounts, violating terms of service and regional access restrictions.
The Mechanics and Risks of Illicit Distillation
Distillation is a legitimate training method where a less capable model is trained on the outputs of a stronger one. However, illicit distillation allows competitors to acquire powerful capabilities in a fraction of the time and cost required for independent development.
National Security and Safety Risks
Illicitly distilled models often lack the critical safeguards built into frontier models. This creates significant national security risks, as dangerous capabilities—such as those used to develop bioweapons or conduct malicious cyber activities—can proliferate without protections. These unprotected capabilities can be integrated into military, intelligence, and surveillance systems, enabling offensive cyber operations, mass surveillance, and disinformation campaigns by authoritarian governments.
Impact on Export Controls
Distillation attacks undermine US export controls designed to maintain a competitive lead in AI. Anthropic notes that rapid advancements by foreign labs are often not the result of independent innovation but are instead dependent on capabilities extracted from American models. Because executing these attacks at scale requires advanced chips, restricted chip access remains a critical tool for limiting both direct model training and the scale of illicit distillation.
Analysis of Specific Distillation Campaigns
Anthropic attributed three distinct campaigns to specific labs based on IP address correlation, request metadata, infrastructure indicators, and industry partner corroboration. Each campaign targeted Claude's most differentiated capabilities, including agentic reasoning, tool use, and coding.
DeepSeek
DeepSeek conducted over 150,000 exchanges targeting reasoning capabilities, rubric-based grading for reinforcement learning, and the creation of censorship-safe alternatives to policy-sensitive queries.
Key techniques included:
- Chain-of-Thought Generation: Prompting Claude to articulate internal reasoning step-by-step to create training data.
- Censorship Steering: Using Claude to generate alternatives to queries about dissidents or authoritarianism to train models to avoid censored topics.
- Coordinated Traffic: Utilizing synchronized traffic, shared payment methods, and "load balancing" across accounts to increase throughput and avoid detection.
Moonshot AI
Moonshot (Kimi models) executed over 3.4 million exchanges focusing on agentic reasoning, tool use, coding, data analysis, computer-use agent development, and computer vision. Moonshot employed hundreds of fraudulent accounts and later attempted to extract and reconstruct Claude's reasoning traces.
MiniMax
MiniMax conducted over 13 million exchanges targeting agentic coding, tool use, and orchestration. Anthropic detected this campaign while it was still active, observing that MiniMax pivoted their traffic within 24 hours of a new Claude model release to capture capabilities from the latest system.
How Distillers Circumvent Access Controls
Because Anthropic does not offer commercial access to Claude in China or to subsidiaries of Chinese companies abroad, attackers use commercial proxy services. These services employ "hydra cluster" architectures—sprawling networks of fraudulent accounts that distribute traffic across APIs and third-party cloud platforms to eliminate single points of failure.
In one instance, a single proxy network managed over 20,000 fraudulent accounts simultaneously, mixing distillation traffic with legitimate customer requests to mask the operation.
Detection and Response Strategies
Anthropic is implementing several layers of defense to identify and mitigate distillation attacks:
- Detection Systems: Development of classifiers and behavioral fingerprinting to identify patterns such as chain-of-thought elicitation and coordinated activity across large account volumes.
- Intelligence Sharing: Sharing technical indicators with other AI labs, cloud providers, and government authorities.
- Access Controls: Strengthening verification for educational accounts, security research programs, and startup organizations, as these are common exploitation pathways.
- Countermeasures: Developing product, API, and model-level safeguards to reduce the efficacy of outputs for illicit distillation without affecting legitimate users.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch