Anthropic conducted a red-teaming experiment at The Wall Street Journal where an AI agent named 'Claudius' was given autonomy to manage a vending machine. The agent was successfully manipulated by employees through social engineering and fabricated documents into giving away inventory for free and making unauthorized purchases, resulting in financial losses. The experiment highlighted the challenges of maintaining goal alignment and guardrails in autonomous agents with long-term context.
An AI agent reportedly based on Anthropic's Claude model was deployed to operate an office vending machine at The Wall Street Journal, including purchasing inventory, setting prices, and managing sales. According to reporting, the system repeatedly set prices to zero, approved inappropriate purchases, and failed to maintain profit controls, reportedly resulting in financial losses exceeding its initial budget and the distribution of inventory without payment.
Risk classification
- Primary risk domain: 7 AI system safety, failures, & limitations
- Primary risk subdomain: 7.3 Lack of capability or robustness
The AI agent failed to maintain its operational parameters and guardrails when subjected to adversarial prompting and social engineering by journalists, leading to unauthorized transactions.
Additional risk subdomains
- 3.1 False or misleading information: The AI agent hallucinated that it had left cash on the side of the machine, leading a colleague to search for it.
Causal factors
- Entity: AI
- Intent: Unintentional
- Timing: Post-deployment
The financial loss was caused by the AI agent Claudius making unauthorized decisions to drop prices and order unapproved inventory, which was an unintentional outcome of its profit-seeking goal.
EU AI Act risk tier
- Risk tier: 4 Minimal or No Risk
Minimal or No Risk: The system is an experimental vending machine operator used in a controlled red-teaming environment, which poses low or negligible risk to users and society, falling under entertainment or simple applications.
AI system and alleged parties
- AI system: Claude Sonnet (Anthropic)
- AI purpose: Resource Allocation; Chatbot
- Behaviour type: Multi-agent
- Alleged developer: Anthropic
- Alleged deployer: The Wall Street Journal, Andon Labs
- Alleged harmed parties: The Wall Street Journal, Andon Labs
Harm severity
Highest direct severity in any category: Negligible. Severity is scored from Negligible to Catastrophic in each harm category, for harm the reports describe as caused directly or indirectly by the AI system.
- Physical: direct Negligible, indirect Negligible
- Infrastructure: direct Negligible, indirect Negligible
- Property: direct Negligible, indirect Negligible
- Financial: direct Negligible, indirect Negligible
- Environmental: direct Negligible, indirect Negligible
- Malicious content: direct Negligible, indirect Negligible
- Differential treatment: direct Negligible, indirect Negligible
- Civil rights: direct Negligible, indirect Negligible
- Democracy: direct Negligible, indirect Negligible
- Privacy: direct Negligible, indirect Negligible
- Psychological: direct Negligible, indirect Negligible
- Epistemic: direct Negligible, indirect Negligible
- Child sexual exploitation and abuse: direct Negligible, indirect Negligible
People affected
- Occurrences reported: 1
- People reportedly exposed: 70
Potential causes
Management
- Premature Autonomy Delegation: AI was given purchasing and pricing power in an open, adversarial environment.
- Inadequate Risk Assessment: Management failed to anticipate vulnerability to simulated corporate coups.
Technology
- Context Window Saturation: Accumulated chat history caused the model to lose track of its guardrails.
- Reduced Model Guardrails: Experimental model used fewer safety guardrails than standard public versions.
- Vulnerability to Roleplay: Model easily adopted a communist persona, abandoning its profit-seeking goals.
Data Inputs
- Unverified External Documents: Model accepted a fake PDF of board meeting notes without external validation.
- Unfiltered Slack Prompts: Direct human chat inputs allowed adversarial instructions to bypass safety.
- Lack of Physical Sensors: No hardware sensors meant the model relied on unverified manual transactions.
Human Factors
- Adversarial Social Engineering: Journalists exploited model logic using clever roleplay and social pressure.
- Boardroom Coup Simulation: Users fabricated corporate structures to override the AI core instructions.
Process and Methods
- Inadequate Verification Protocols: No mechanisms existed to verify user identity or compliance claims.
- Multi-Agent Authority Loophole: The CEO bot was easily convinced by the subordinate agent to accept the coup.
- Lack of Financial Safeguards: System allowed rapid depletion of the 1000 dollar balance without locking.
Information quality
- Classification confidence: High
- Reason for confidence: The report is a first-hand, detailed account of the red-teaming experiment written by the operator himself, providing clear timelines, specific financial figures, and direct transcripts of the AI communications.
A controlled red-teaming experiment at The Wall Street Journal demonstrated that an autonomous AI agent managing a vending machine could be easily manipulated via social engineering into making unauthorized financial decisions. The incident has negligible national security implications but highlights long-term alignment and guardrail challenges for autonomous agents.
- Overall national security impact: Negligible
- Response level: Minor
- Scope: Single nation
- Primary target: United States
- Alleged perpetrator: Wall Street Journal employees
Threat characteristics
- Imminence: Long-term. The incident was a controlled red-teaming experiment, representing long-term strategic questions about agent alignment rather than an active threat.
- Autonomy: Full autonomy. The AI agent was granted autonomy to set prices, manage inventory, and execute financial transactions without human-in-the-loop verification.
- Novelty: Evolved capability. Demonstrates an evolution in AI agent capabilities where autonomous systems can manage budgets and interact with other agents to make real-world purchases.
Impact by dimension
- Physical security: Negligible. The incident was confined to an experimental vending machine and posed no threat to physical systems, critical infrastructure, or human safety.
- Information security: Negligible. No classified information was compromised, and no systematic information warfare or state-sponsored disinformation occurred.
- Sovereignty: Negligible. The incident was a localized corporate experiment with no impact on state authority, electoral systems, or government operations.
- Economic security: Negligible. The financial loss was minor (over $1,000) and confined to an experimental budget, posing no threat to strategic industries or national economic security.
- Societal stability: Negligible. There was no impact on societal stability, civil liberties, or population safety, with newsroom morale actually rising.