Players Manipulated GPT-3-Powered Game to Generate Sexually Explicit Material Involving Children

The AI-powered game AI Dungeon, which utilized OpenAI's GPT-3, was found to generate sexually explicit content involving children when prompted by users. This led to a conflict between the developer, Latitude, and OpenAI regarding content moderation. Additionally, a security vulnerability in the game resulted in the exposure of hundreds of thousands of private user-generated stories to the public.

Latitude's GPT-3-powered game AI Dungeon was reportedly abused by some players who manipulated its AI to generate sexually explicit stories involving children.

Source: AI Incident Database

Risk classification

  • Primary risk domain: 1 Discrimination & Toxicity
  • Primary risk subdomain: 1.2 Exposure to toxic content

The GPT-3 model generated stories depicting sexual encounters involving children, which constitutes exposure to toxic content and child sexual abuse material.

Additional risk subdomains

  • 2.1 Compromise of privacy by obtaining, leaking or correctly inferring sensitive information: A security vulnerability exposed hundreds of thousands of private, user-generated stories to the public.

Causal factors

  • Entity: AI
  • Intent: Unintentional
  • Timing: Post-deployment

The generation of sexually explicit content involving children was an unexpected and unintentional output of the GPT-3 model post-deployment, triggered by user prompts.

EU AI Act risk tier

  • Risk tier: 3 Limited Risk

Limited Risk: The system is a generative AI text creator and interactive game, which falls under transparency obligations for AI-generated content.

AI system and alleged parties

  • AI system: GPT-3 3 (OpenAI)
  • AI purpose: Game Content Generation; Writing Assistant
  • Behaviour type: Assistant
  • Alleged developer: OpenAI, Latitude
  • Alleged deployer: Latitude
  • Alleged harmed parties: Latitude

Harm severity

Highest direct severity in any category: Substantial. Severity is scored from Negligible to Catastrophic in each harm category, for harm the reports describe as caused directly or indirectly by the AI system.

  • Physical: direct Negligible, indirect Negligible
  • Infrastructure: direct Negligible, indirect Negligible
  • Property: direct Negligible, indirect Negligible
  • Financial: direct Negligible, indirect Negligible
  • Environmental: direct Negligible, indirect Negligible
  • Malicious content: direct Minor, indirect Negligible
  • Differential treatment: direct Negligible, indirect Negligible
  • Civil rights: direct Negligible, indirect Negligible
  • Democracy: direct Negligible, indirect Negligible
  • Privacy: direct Substantial, indirect Minor
  • Psychological: direct Negligible, indirect Minor
  • Epistemic: direct Negligible, indirect Negligible
  • Child sexual exploitation and abuse: direct Minor, indirect Negligible

Malicious content

Reported: The report explicitly describes the generation of toxic and sexually explicit content involving children.

Directly caused: The AI system generated stories depicting sexual encounters involving children when prompted by some players.

Indirectly caused: N/A

Inferred additional harm: It is likely that many more instances of toxic or sexually explicit content involving minors were generated than the sampled stories analyzed by the user.

Privacy

Reported: The report explicitly describes privacy violations due to a security flaw and Latitude's content monitoring policies.

Directly caused: A security flaw made every story generated in the game publicly accessible, allowing a user to download several hundred thousand private adventures.

Indirectly caused: Latitude's plans to manually review flagged content led to concerns that staff would snoop on private, fictional creations of users.

Inferred additional harm: The exposure of hundreds of thousands of private stories likely compromised the privacy of thousands of active players who expected their fictional writings to remain confidential.

Psychological

Reported: The report explicitly describes psychological distress and discomfort among users.

Directly caused: Several players sent examples of unbidden sexual themes generated by the AI that left them 'feeling deeply uncomfortable'.

Indirectly caused: Users expressed feeling 'betrayed' and angry over Latitude's sudden implementation of invasive content moderation and manual reviews of their private stories.

Inferred additional harm: N/A

Child sexual exploitation and abuse

Reported: The report explicitly describes a CSEA incident involving AI-generated stories depicting sexual encounters with children.

Directly caused: The AI system generated stories depicting sexual encounters involving children, which was flagged by OpenAI's monitoring system.

Indirectly caused: N/A

Inferred additional harm: N/A

People affected

  • Occurrences reported: 1
  • People reportedly harmed: 100
  • People reportedly exposed: 20000

Potential causes

Management

  • Prioritizing Growth Over Safety: Latitude offered unconstrained access to GPT-3 without robust guardrails.
  • Upstream Pressure from OpenAI: OpenAI demanded immediate action after discovering policy violations.

Technology

  • Unpredictable Text Generation: GPT models generate text based on patterns without semantic understanding.
  • Oversensitive Algorithmic Filters: Filters blocked benign phrases like '8-year-old laptop' causing user anger.

Data Inputs

  • Unfiltered Training Data: Models trained on internet text containing toxic and inappropriate content.
  • User Prompt Manipulation: Users input creative prompts to steer the AI into generating NSFW stories.

Human Factors

  • User Backlash Over Privacy: Players revolted against manual reviews of their private fictional stories.
  • Intentional Policy Evasion: Some players actively tried to generate prohibited child sexual content.

Process and Methods

  • Inadequate Content Moderation: Latitude relied on basic word filters that were easily bypassed or too strict.
  • Intrusive Manual Review Plan: Reading flagged private stories violated user trust and privacy expectations.

Regulatory Environment

  • Lack of AI Safety Standards: No industry-wide regulations governed the deployment of generative models.

Information quality

  • Classification confidence: High
  • Reason for confidence: The report provides clear, detailed accounts of the AI system used (GPT-3), the developer (Latitude), the specific safety failures (generation of child sexual content and a security vulnerability leaking stories), and the resulting user backlash. The facts are well-documented with quotes from both OpenAI's CEO and Latitude's spokespersons.
  • Ambiguities identified: The exact number of users whose private stories were accessed by unauthorized parties during the security flaw is not fully specified, though a sample of 188,000 was analyzed.
  • Alternative interpretations: The incident could be viewed primarily as a software security failure rather than an AI safety failure, though the core conflict stemmed from the AI's generation of toxic content.

An early commercial LLM deployment (AI Dungeon) was abused by users to generate child-exploitation-adjacent text, coupled with a security vulnerability leaking user data. While presenting serious ethical and privacy concerns, the national security impact is negligible to minor, with no threat to physical infrastructure, intelligence, or state sovereignty.

  • Overall national security impact: Minor
  • Response level: Moderate
  • Scope: Multiple nations
  • Primary target: No clear primary
  • Other affected: Unknown
  • Alleged perpetrator: Individual users

Threat characteristics

  • Imminence: Long-term. The immediate software vulnerability and content moderation issues were resolved, leaving long-term policy and safety lessons.
  • Autonomy: Human-controlled. The AI acted as a co-writing assistant, generating content directly in response to and guided by human user prompts.
  • Novelty: Evolved capability. Demonstrates a significant advancement in the scale and realism of AI-generated toxic content and associated privacy risks in commercial LLM deployments.

Impact by dimension

  • Physical security: Negligible. No threat to physical systems, critical infrastructure, or human safety was identified in this text-based gaming incident.
  • Information security: Negligible. No compromise of classified intelligence or state-sponsored disinformation campaigns occurred; the leaked data consisted of fictional game stories.
  • Sovereignty: Negligible. The incident did not impact state authority, electoral systems, or core government operations.
  • Economic security: Negligible. Impact was limited to commercial and reputational damage to a single startup company, posing no threat to national economic security.
  • Societal stability: Minor. The generation of child-exploitation-adjacent text and a mass privacy leak of user stories affected individual users, but did not threaten national-scale societal stability.
Explore in the interactive Incident Tracker