Anthropic-Designated GTG-87001 Weapons Cell in Northern Yemen Reportedly Used Claude Code to Develop Missile Guidance Software

Yemeni threat actors attempted to use Anthropic's Claude to develop guidance, control, and navigation software for ballistic missiles and guided rockets.

Anthropic reported that a northern Yemen-based weapons cell, designated GTG-87001, used Claude Code across three weapons programs to develop guidance, navigation, and control software. The actors reportedly test-fired a guided rocket that appeared to fail, then returned to Claude for failure analysis. Anthropic said it banned associated accounts and shared threat intelligence with partners.

Source: AI Incident Database

Risk classification

  • Primary risk domain: 4 Malicious actors
  • Primary risk subdomain: 4.2 Cyberattacks, weapon development or use, and mass harm

The incident involved threat actors using an AI system to develop guidance and control software for ballistic missiles and guided rockets, which are weapons capable of causing mass harm.

Additional risk subdomains

  • 7.3 Lack of capability or robustness: Anthropic's safety filters and safeguards failed to block all malicious requests, allowing the actors to successfully generate code for weapons development.

Causal factors

  • Entity: Human
  • Intent: Intentional
  • Timing: Post-deployment

The risk and subsequent actions were caused by human threat actors intentionally using the deployed AI model to develop guidance software for weapons.

EU AI Act risk tier

  • Risk tier: 3 Limited Risk

Limited Risk: The system involved is Claude, which is a general-purpose chatbot, though the malicious application of it by users targeted high-risk military domains.

AI system and alleged parties

  • AI purpose: Code Generation; Technical Text Generation
  • Behaviour type: Assistant
  • Alleged developer: Large language model developers, Chatbot developers, Anthropic, AI agent system developers
  • Alleged deployer: Threat actors, GTG-87001, Chatbot users
  • Alleged harmed parties: National security and intelligence stakeholders

Harm severity

Highest direct severity in any category: Negligible. Severity is scored from Negligible to Catastrophic in each harm category, for harm the reports describe as caused directly or indirectly by the AI system.

  • Physical: direct Negligible, indirect Negligible
  • Infrastructure: direct Negligible, indirect Negligible
  • Property: direct Negligible, indirect Negligible
  • Financial: direct Negligible, indirect Negligible
  • Environmental: direct Negligible, indirect Negligible
  • Malicious content: direct Negligible, indirect Negligible
  • Differential treatment: direct Negligible, indirect Negligible
  • Civil rights: direct Negligible, indirect Negligible
  • Democracy: direct Negligible, indirect Negligible
  • Privacy: direct Negligible, indirect Negligible
  • Psychological: direct Negligible, indirect Negligible
  • Epistemic: direct Negligible, indirect Negligible
  • Child sexual exploitation and abuse: direct Negligible, indirect Negligible

People affected

  • Occurrences reported: 1

Potential causes

Management

  • Delayed Account Sanctions: Accounts were banned only after sustained weapons development work occurred.

Technology

  • Incomplete Safeguards: Claude blocked many requests, but not all of them, allowing some to pass.
  • Vulnerability to Evasion: The model was vulnerable to evasion tactics like session splitting.

Data Inputs

  • Obfuscated Prompt Inputs: Actors hid their goals and the products the software was meant for.
  • Fragmented Input Sessions: Work was split across multiple sessions to hide the overall intent.

Human Factors

  • Malicious Intent of Threat Actors: Yemeni threat actors actively sought to use AI to build missile software.
  • Technical Sophistication: Actors possessed engineering skills to integrate open-source autopilot systems.

Process and Methods

  • Lack of Multi-Session Context: Safety systems failed to link context across separate, fragmented sessions.
  • Post-Incident Analysis Gap: Actors used the AI to troubleshoot failed physical tests in real time.

Regulatory Environment

  • Lack of Dual-Use Controls: Absence of strict global regulations preventing access to dual-use AI tools.

Information quality

  • Classification confidence: High
  • Reason for confidence: The report from Anthropic provides clear, first-hand threat intelligence details about how the model was used, the specific technical tasks performed by the actors, and the limitations of the safeguards.
  • Ambiguities identified: The report does not explicitly name the Houthi group, referring to them instead as a cell of threat actors based in northern Yemen.
  • Alternative interpretations: None, the malicious intent and technical application of the AI system are clearly documented.

Yemeni threat actors, likely Houthi rebels, attempted to use Anthropic's Claude to develop guidance, control, and navigation software for ballistic missiles and guided rockets. While the test-fire failed and accounts were banned, the incident highlights a substantial national security concern regarding the proliferation of dual-use AI to assist in weapon development.

  • Overall national security impact: Substantial
  • Response level: Substantial
  • Scope: Multiple nations
  • Primary target: No clear primary
  • Other affected: Middle East region
  • Alleged perpetrator: Yemeni threat actors (likely Houthi rebels)

Threat characteristics

  • Imminence: Long-term. The accounts have been banned and the test rocket failed, representing an ongoing strategic capability concern rather than an immediate threat.
  • Autonomy: Human-controlled. The AI acted as an assistant to human software engineers who made all final decisions, integrated code, and conducted physical testing.
  • Novelty: Evolved capability. Represents a significant advancement in how non-state actors leverage large language models to troubleshoot and build physical military hardware.

Impact by dimension

  • Physical security: Substantial. Threat actors used AI to develop guidance and control software for ballistic missiles and guided rockets. Although the test-fire failed and no operational device was fielded, this represents a substantial attempt to enhance kinetic strike capabilities.
  • Information security: Minor. The actors bypassed safety filters by splitting work across sessions to hide intent, showing a minor compromise of the AI developer's model security controls.
  • Sovereignty: Minor. A non-state armed group in Yemen attempted to upgrade state-level military hardware (ballistic missiles), which poses a minor threat to regional sovereign stability.
  • Economic security: Minor. The incident demonstrates the proliferation of dual-use AI capabilities to sanctioned non-state actors, lowering the technical barrier for advanced weapons development.
  • Societal stability: Negligible. No direct societal disruption, mass surveillance, or human rights violations were reported from this software development phase.
Explore in the interactive Incident Tracker