Anthropic reported several instances of malicious actors misusing their LLM in March 2025. These included an influence-as-a-service operation orchestrating over 100 social media bots, a credential stuffing campaign targeting IoT security cameras, recruitment fraud in Eastern Europe, and a novice actor developing sophisticated malware. Anthropic identified and banned the accounts involved in these activities.
In April 2025, Anthropic published a report detailing several misuse cases involving its Claude LLM, all detected in March. These included an "influence-as-a-service" operation that orchestrated over 100 social media bots; an effort to scrape and test leaked credentials for security camera access; a recruitment fraud campaign targeting Eastern Europe; and a novice actor developing sophisticated malware. Anthropic banned the accounts involved but could not confirm downstream deployment.
Risk classification
- Primary risk domain: 4 Malicious actors
- Primary risk subdomain: 4.1 Disinformation, surveillance, and influence at scale
The primary and most novel misuse described is a coordinated influence-as-a-service campaign using the AI to orchestrate over 100 social media bots to manipulate public opinion and political narratives at scale.
Additional risk subdomains
- 4.2 Cyberattacks, weapon development or use, and mass harm: The AI was used by threat actors to develop advanced malware and build brute-force tools targeting IoT security cameras.
- 4.3 Fraud, scams, and targeted manipulation: The AI was used to refine and sanitize recruitment scam messages targeting Eastern European job seekers to make them more convincing.
Causal factors
- Entity: Human
- Intent: Intentional
- Timing: Post-deployment
The incident was caused by human threat actors intentionally misusing the deployed AI model for malicious purposes post-deployment.
EU AI Act risk tier
- Risk tier: 1 Unacceptable
Unacceptable Risk: The report describes the AI being used for an 'influence-as-a-service' operation to orchestrate social media bots and make tactical decisions to manipulate public opinion, which constitutes 'AI used for subliminal manipulation or behavioral influence leading to harm'.
AI system and alleged parties
- AI system: Claude (Anthropic)
- AI purpose: Social Media Content Generation; Chatbot
- Behaviour type: Agent
- Alleged developer: Anthropic
- Alleged deployer: Unknown malicious actors, Unknown cybercriminals, Influence-as-a-service operators
- Alleged harmed parties: social media users, People targeted by malware, National security and intelligence stakeholders, Job seekers in Eastern Europe, IoT security camera owners, Epistemic integrity
Harm severity
Highest direct severity in any category: Substantial. Severity is scored from Negligible to Catastrophic in each harm category, for harm the reports describe as caused directly or indirectly by the AI system.
- Physical: direct Negligible, indirect Negligible
- Infrastructure: direct Negligible, indirect Negligible
- Property: direct Negligible, indirect Negligible
- Financial: direct Negligible, indirect Negligible
- Environmental: direct Negligible, indirect Negligible
- Malicious content: direct Minor, indirect Minor
- Differential treatment: direct Negligible, indirect Negligible
- Civil rights: direct Negligible, indirect Negligible
- Democracy: direct Minor, indirect Minor
- Privacy: direct Minor, indirect Negligible
- Psychological: direct Negligible, indirect Negligible
- Epistemic: direct Minor, indirect Substantial
- Child sexual exploitation and abuse: direct Negligible, indirect Negligible
Malicious content
Reported: Yes
Directly caused: The AI directly generated politically-aligned social media posts and responses in multiple languages for over 100 bot accounts, as well as polished recruitment scam messages.
Indirectly caused: The generated content was used to populate a network of bot accounts that engaged with tens of thousands of authentic social media users.
Inferred additional harm: N/A
Democracy
Reported: Yes
Directly caused: The AI orchestrated bot accounts to promote specific political narratives supporting or undermining European, Iranian, UAE, and Kenyan interests.
Indirectly caused: This coordinated activity mimics state-affiliated campaigns, potentially distorting public discourse and political processes across multiple countries.
Inferred additional harm: N/A
Privacy
Reported: Yes
Directly caused: The AI was used to process information stealer logs from Telegram and enhance tools to scrape leaked credentials associated with IoT security cameras.
Indirectly caused: N/A
Inferred additional harm: N/A
Epistemic
Reported: Yes
Directly caused: The AI generated consistent, politically-aligned personas and responses to deceive authentic users into believing they were interacting with real humans.
Indirectly caused: The operation engaged with tens of thousands of authentic accounts, polluting the social media information ecosystem with synthetic engagement.
Inferred additional harm: N/A
People affected
- Occurrences reported: 4
- People reportedly exposed: 20000
Potential causes
Management
- Over-Reliance on Post-Hoc Bans: Reactive account banning occurred after the malicious activity took place.
- Insufficient Threat Intel Sharing: Anthropic's public report lacked actionable intelligence like prompts or IOCs.
- Inadequate Pre-Release Testing: Model safety guardrails did not fully anticipate agentic orchestration risks.
Technology
- Dual-Use Model Capabilities: Model's advanced coding and reasoning can be repurposed for malicious acts.
- Agentic Orchestration Capabilities: Model can make tactical decisions to orchestrate networks of social bots.
- Natural Language Generation: Generates highly convincing, multilingual text that mimics human humor.
Data Inputs
- Malicious Prompting: Adversaries input instructions to bypass safety and generate harmful outputs.
- Structured JSON Data Formats: Structured inputs and outputs allowed efficient scaling of persona management.
- Integration of Stealer Logs: Processing raw Telegram logs and breach data to target IoT devices.
Human Factors
- Novice Actor Upskilling: Low-skilled actors use AI to flatten the learning curve for malware.
- Convincing Social Engineering: Actors leverage polished language to dupe unsuspecting job seekers.
- Adversarial Intent: Malicious actors actively seek to exploit AI models for financial gain.
Process and Methods
- Lack of Real-Time Prompt Detection: Standard safety classifiers failed to block all adversarial prompt patterns.
- Absence of Shared Indicators: Lack of shared IOCs, IP addresses, or prompts hinders industry defense.
- Inadequate Post-Deployment Auditing: Abuse was detected after operations had already scaled and engaged users.
Regulatory Environment
- Weak Social Media Bot Controls: Social media platforms failed to prevent automated bot networks.
- Lack of AI Safety Standards: Absence of standardized frameworks for evaluating AI-driven influence ops.
- Cross-Border Cyber Regulation Gaps: Difficult to regulate commercial influence services operating across borders.
Information quality
- Classification confidence: High
- Reason for confidence: The reports are based on direct threat intelligence findings from Anthropic, detailing specific case studies of misuse of their model. While some technical details like exact prompts or indicators of compromise are omitted for security, the descriptions of the activities, their scale, and the mitigation actions taken are clear and authoritative.
- Ambiguities identified: The exact identity of the threat actors and their clients remains unknown, and the real-world success or deployment of the malware, credential stuffing, and recruitment fraud campaigns could not be fully confirmed.
- Alternative interpretations: None. The activities are clearly classified as malicious misuses of a deployed LLM.
Anthropic disrupted several malicious campaigns misusing its Claude LLM, most notably an influence-as-a-service operation where the AI autonomously orchestrated over 100 social media bots to influence political narratives. Other misuses included credential stuffing, Eastern European recruitment fraud, and a novice actor developing advanced malware. While the direct real-world impacts were mitigated, these incidents demonstrate how frontier AI models are evolving to lower technical barriers for complex cyber and information operations.
- Overall national security impact: Substantial
- Response level: Substantial
- Scope: Multiple nations
- Primary target: No clear primary
- Other affected: Europe, Iran, UAE, Kenya, Eastern Europe
- Alleged perpetrator: Unknown
Threat characteristics
- Imminence: Long-term. The specific campaigns were identified and disrupted by the provider, representing an ongoing strategic capability concern rather than an active, unmitigated crisis.
- Autonomy: Human-supervised. The AI acted autonomously in generating content and making tactical decisions on bot interactions, but operated under the strategic direction and setup of human threat actors.
- Novelty: Evolved capability. Represents a significant advancement in existing threats, showing LLMs transitioning from simple content generation to active, tactical orchestration of botnets and lowering barriers for novice malware developers.
Impact by dimension
- Physical security: Minor. Malicious actors targeted IoT security cameras with brute-force tools and developed malware with facial recognition, representing minor physical and surveillance security risks.
- Information security: Substantial. An influence-as-a-service operation used AI to orchestrate over 100 social media bots making autonomous tactical engagement decisions, interacting with tens of thousands of users to push geopolitical narratives.
- Sovereignty: Minor. The botnet operation targeted narratives surrounding European, Iranian, UAE, and Kenyan interests, representing minor foreign interference in sovereign political discourse.
- Economic security: Minor. The AI was used for recruitment fraud in Eastern Europe and credential harvesting, presenting minor threat to economic security and individual financial safety.
- Societal stability: Minor. Coordinated botnets and recruitment scams targeted vulnerable populations (job seekers, social media users) but did not lead to large-scale civil unrest or systemic rights violations.