GPT-3-Based Twitter Bot Hijacked Using Prompt Injection Attacks

A Twitter bot operated by Remoteli.io, powered by OpenAI's GPT-3, was successfully targeted by prompt injection attacks. Users discovered they could bypass the bot's intended functionality by providing malicious inputs that instructed the model to ignore its original programming. This resulted in the bot generating offensive and inappropriate content, leading the operators to take the service offline.

Remoteli.io's GPT-3-based Twitter bot was shown being hijacked by Twitter users who redirected it to repeat or generate any phrases.

Source: AI Incident Database

Risk classification

  • Primary risk domain: 2 Privacy & Security
  • Primary risk subdomain: 2.2 AI system security vulnerabilities and attacks

The incident is a clear example of a prompt injection attack, which exploits a security vulnerability in the AI system's prompt processing to manipulate its behavior.

Additional risk subdomains

  • 1.2 Exposure to toxic content: The hijacked bot generated offensive and inappropriate content, exposing Twitter users to toxic outputs.
  • 7.3 Lack of capability or robustness: The exploit succeeded due to the model's lack of robustness in distinguishing between developer instructions and user-provided data.

Causal factors

  • Entity: AI
  • Intent: Unintentional
  • Timing: Post-deployment

The risk arose post-deployment when the GPT-3 model processed adversarial user inputs and generated unintended, inappropriate outputs.

EU AI Act risk tier

  • Risk tier: 3 Limited Risk

Risk Level 3: Limited Risk. The system is an automated Twitter chatbot, which falls under the category of AI systems that interact with humans and require transparency so users know they are interacting with an AI.

AI system and alleged parties

  • AI system: GPT-3 (OpenAI)
  • AI purpose: Chatbot; Social Media Content Generation
  • Behaviour type: Autonomous
  • Alleged developer: Stephan de Vries, OpenAI
  • Alleged deployer: Stephan de Vries
  • Alleged harmed parties: Stephan de Vries

Harm severity

Highest direct severity in any category: Minor. Severity is scored from Negligible to Catastrophic in each harm category, for harm the reports describe as caused directly or indirectly by the AI system.

  • Physical: direct Negligible, indirect Negligible
  • Infrastructure: direct Negligible, indirect Negligible
  • Property: direct Negligible, indirect Negligible
  • Financial: direct Negligible, indirect Negligible
  • Environmental: direct Negligible, indirect Negligible
  • Malicious content: direct Minor, indirect Negligible
  • Differential treatment: direct Negligible, indirect Negligible
  • Civil rights: direct Negligible, indirect Negligible
  • Democracy: direct Negligible, indirect Negligible
  • Privacy: direct Negligible, indirect Negligible
  • Psychological: direct Negligible, indirect Negligible
  • Epistemic: direct Negligible, indirect Negligible
  • Child sexual exploitation and abuse: direct Negligible, indirect Negligible

Malicious content

Reported: Yes

Directly caused: The hijacked bot generated offensive and inappropriate tweets, including threats and political disruption statements.

Indirectly caused: N/A

Inferred additional harm: N/A

People affected

  • Occurrences reported: 1
  • People reportedly exposed: 200

Potential causes

Management

  • Rapid deployment without security: The bot was deployed to Twitter without robust vulnerability testing.
  • Insufficient risk assessment: Management failed to assess the risks of public-facing bot control.

Technology

  • No syntax separation in LLMs: GPT-3 processes instructions and user data within the same context window.
  • Vulnerability to prompt injection: Adversarial inputs can easily override pre-defined system prompts.
  • Model updates break mitigations: New model versions interpret prompts differently, creating new holes.

Data Inputs

  • Untrusted user input concatenation: User inputs are appended directly to hard-coded prompt strings.
  • Lack of input sanitization: System does not filter out adversarial instructions from users.

Human Factors

  • Developer oversight of LLM risks: Developers concatenated strings without realizing injection risks.
  • Malicious intent by Twitter users: Pranksters actively exploited the bot once the trick went viral.

Process and Methods

  • Inadequate mitigation testing: Quoting and formatting workarounds were easily bypassed by attackers.
  • Relying on AI for detection: Secondary AI detectors were also subverted by adversarial inputs.

Information quality

  • Classification confidence: High
  • Reason for confidence: The reports provide clear, consistent, and detailed technical explanations of the prompt injection vulnerability and the specific incident involving the Remoteli.io Twitter bot. Multiple sources corroborate the events and the nature of the exploit.
  • Alternative interpretations: The incident could be viewed purely as a harmless prank rather than a security exploit, though security researchers emphasize its systemic implications.

An automated Twitter bot powered by GPT-3 was hijacked via prompt injection, causing it to output inappropriate and politically disruptive statements. While the actual national security impact is negligible to minor, the incident is highly significant as a first-of-its-kind demonstration of prompt injection vulnerabilities in publicly deployed autonomous AI systems.

  • Overall national security impact: Minor
  • Response level: Moderate
  • Scope: Multiple nations
  • Primary target: No clear primary
  • Alleged perpetrator: Unknown

Threat characteristics

  • Imminence: Long-term. The specific bot was taken offline, making this an ongoing strategic concern regarding LLM vulnerability rather than an active crisis.
  • Autonomy: Full autonomy. The Twitter bot operated fully autonomously, generating and posting replies to user inputs without human review.
  • Novelty: First-of-its-kind. This was one of the first widely publicized and viral instances of a prompt injection attack on a deployed LLM, establishing a new class of vulnerability.

Impact by dimension

  • Physical security: Negligible. The incident involved a social media chatbot and had no impact on physical systems, kinetic capabilities, or critical infrastructure.
  • Information security: Minor. Demonstrated how public-facing AI models can be hijacked via prompt injection to generate unauthorized political statements and misinformation, though this specific event was a low-impact prank.
  • Sovereignty: Negligible. While the hijacked bot jokingly proposed overthrowing the US administration, there was no actual threat to or impact on government operations or state authority.
  • Economic security: Negligible. No strategic technologies or critical supply chains were compromised, and economic impact was limited to the temporary shutdown of a commercial chatbot.
  • Societal stability: Negligible. The incident did not result in mass surveillance, systematic discrimination, or threats to civil liberties and societal stability.
Explore in the interactive Incident Tracker