Detecting and Countering Malicious Uses of Claude

Anthropic has identified and banned several actors using Claude to orchestrate influence operations, automate credential scraping, enhance recruitment fraud, and develop malware. This effort highlights an emerging trend where frontier AI models are used not only for content generation but as orchestrators for complex, semi-autonomous abuse systems.

Orchestration of Influence-as-a-Service Operations

Anthropic detected a professional "influence-as-a-service" operation that used Claude to manage over 100 social media bot accounts across Facebook and X (formerly Twitter). Unlike traditional AI misuse, this operation used Claude as an orchestrator to make tactical engagement decisions—determining whether bots should like, share, comment on, or ignore specific posts based on politically motivated personas.

Operation Details

  • Scale: The operation engaged with tens of thousands of authentic social media accounts across multiple countries and languages.
  • Tactics: Claude was used to create and maintain consistent political personas, generate aligned responses, and create prompts for image-generation tools.
  • Objective: The actor focused on sustained, long-term engagement to promote moderate political perspectives rather than attempting to achieve viral status.
  • Attribution: While the narratives were consistent with state-affiliated campaigns, Anthropic has not confirmed specific attribution.

AI-Enabled Cyberattack and Fraud Capabilities

Anthropic identified three distinct cases where AI was used to lower the technical barrier for malicious activities or enhance existing attack vectors.

Credential Scraping and IoT Access

A sophisticated actor used Claude to develop tools for scraping leaked usernames and passwords associated with security cameras. The actor integrated commercial breach data platforms and private stealer log communities to identify targets. Claude was used to rewrite open-source scraping toolkits, create scripts for target URL extraction, and improve backend search features. Anthropic has not confirmed if these capabilities were successfully deployed.

Recruitment Fraud and Language Sanitization

An operation targeting job seekers in Eastern European countries used Claude for "real-time language sanitization." By submitting poorly written, non-native English text and asking Claude to rewrite it as a native speaker, the actors improved the perceived legitimacy of their scams. Claude was also used to develop convincing recruitment narratives and interview scenarios. Anthropic has not confirmed successful deployment of these scams.

Malware Development for Novice Actors

Anthropic observed a novice actor with limited coding skills use Claude to rapidly develop sophisticated malicious tools. The actor's toolkit evolved from basic scripts to an advanced suite featuring facial recognition, dark web scanning, and a graphical user interface (GUI) for generating undetectable malicious payloads designed to evade security controls. Anthropic has not confirmed real-world deployment of this malware.

Detection and Mitigation Framework

Anthropic employs a multi-layered intelligence program to identify harms that bypass standard scaled detection. The company utilizes a combination of classifiers—which evaluate user inputs and model responses—and advanced analysis techniques.

Technical Analysis Methods

To analyze large volumes of conversation data and identify patterns of misuse, Anthropic applied techniques from its research, specifically:

  • Clio: A tool for detecting patterns of misuse.
  • Hierarchical Summarization: A method for monitoring and analyzing data at scale.

Upon detecting these activities, Anthropic banned the associated accounts and integrated the learnings from these cases into its broader set of controls to improve future detection and prevention.

Sources

Related

  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch