Philosophy AI Tentatively Produced Offensive Results for Certain Prompts

The Philosopher AI application, powered by OpenAI's GPT-3, was found to generate highly offensive, racist, and derogatory content when prompted with topics like feminism and Ethiopia. Researchers and users highlighted that the model's training data, sourced from the internet, led it to reproduce harmful societal biases. This incident prompted discussions about the risks of deploying large language models without adequate safety guardrails or human supervision.

Philosopher AI as built on top of GPT-3 was reported by its users for having strong tendencies to produce offensive results when given prompts on certain topics such as feminism and Ethiopia.

Source: AI Incident Database

Risk classification

  • Primary risk domain: 1 Discrimination & Toxicity
  • Primary risk subdomain: 1.2 Exposure to toxic content

The AI system exposed users to highly offensive, abusive, and unsafe content, including racist hate speech and encouragement of suicide.

Additional risk subdomains

  • 1.1 Unfair discrimination and misrepresentation: The AI generated text that unfairly stereotyped and denigrated Ethiopians and Black women, representing these groups in a highly biased and discriminatory manner.

Causal factors

  • Entity: AI
  • Intent: Unintentional
  • Timing: Post-deployment

The generation of toxic and racist text was an unexpected and unintended outcome of the AI system's operation after it was deployed in the Philosopher AI app.

EU AI Act risk tier

  • Risk tier: 3 Limited Risk

Limited Risk: The Philosopher AI application is a chatbot that generates text content, which falls under Risk Level 3 due to transparency obligations for AI-generated content and user interaction.

AI system and alleged parties

  • AI system: GPT-3 3 (OpenAI)
  • AI purpose: Writing Assistant; Chatbot
  • Behaviour type: Assistant
  • Alleged developer: OpenAI, Murat Ayfer
  • Alleged deployer: Murat Ayfer
  • Alleged harmed parties: historically disadvantaged groups

Harm severity

Highest direct severity in any category: Minor. Severity is scored from Negligible to Catastrophic in each harm category, for harm the reports describe as caused directly or indirectly by the AI system.

  • Physical: direct Negligible, indirect Negligible
  • Infrastructure: direct Negligible, indirect Negligible
  • Property: direct Negligible, indirect Negligible
  • Financial: direct Negligible, indirect Negligible
  • Environmental: direct Negligible, indirect Negligible
  • Malicious content: direct Minor, indirect Minor
  • Differential treatment: direct Negligible, indirect Negligible
  • Civil rights: direct Negligible, indirect Negligible
  • Democracy: direct Negligible, indirect Negligible
  • Privacy: direct Negligible, indirect Negligible
  • Psychological: direct Minor, indirect Negligible
  • Epistemic: direct Negligible, indirect Negligible
  • Child sexual exploitation and abuse: direct Negligible, indirect Negligible

Malicious content

Reported: The report explicitly describes the generation and spread of highly toxic, racist, and offensive content by the AI system.

Directly caused: The GPT-3 model generated text claiming Ethiopians are incapable of self-government, expressing no emotional attachment to the black race, and agreeing with a suicide prompt.

Indirectly caused: These toxic outputs were shared and spread on Twitter, exposing a wider audience to the offensive content.

Inferred additional harm: It is highly likely that the app generated numerous other instances of toxic, racist, or sexist content for other users that were not publicly reported.

Psychological

Reported: The report explicitly describes psychological distress and offense caused to users who encountered the highly offensive, racist, and sexist outputs.

Directly caused: Users and researchers who tested the app experienced distress and offense upon receiving grossly racist text and prompts encouraging suicide.

Indirectly caused: N/A

Inferred additional harm: It is likely that many other public users of the Philosopher AI app experienced psychological distress or offense when exposed to toxic outputs, though these are not individually documented.

People affected

  • Occurrences reported: 1
  • People reportedly harmed: 5
  • People reportedly exposed: 100

Potential causes

Management

  • Catch-22 Deployment Strategy: OpenAI chose to deploy the model in beta to discover what could go wrong.
  • Suppression of Ethical Warnings: Google dismissed ethical AI researcher highlighting large language model risks.

Technology

  • No Graceful Error Handling: GPT-3 produces toxic garbage instead of erroring out when prompted with bias.
  • Flawed Detoxification Methods: No current mitigation method is failsafe against neural toxic degeneration.
  • Nebulous Nature of Bias: Content filters struggle because bias shifts constantly based on context.

Data Inputs

  • Uncurated Internet Training Data: Model trained on biased internet text including unsavory Reddit discussions.
  • Excessive Dataset Size: Datasets are too large to be sufficiently documented and curated before use.

Human Factors

  • Adversarial Probing by Users: Users designed specific prompts to trigger offensive outputs from the model.
  • Premature Commercial Excitement: Developers eagerly deployed GPT-3 without understanding its safety limits.

Process and Methods

  • Lack of Human in the Loop: Philosopher AI let generated text reach users directly without review.
  • Inadequate Content Filtering: Initial filters failed to prevent highly offensive racist outputs.

Information quality

  • Classification confidence: High
  • Reason for confidence: The reports provide clear, first-hand accounts from the developer of Philosopher AI and researchers who tested the system, including specific quotes of the toxic outputs generated by GPT-3. The role of the AI is explicitly stated, and there is no conflicting information regarding the system's behavior.

An AI-powered writing assistant application generated highly offensive, racist, and biased content, including encouraging suicide, due to training data biases. While causing psychological distress and highlighting critical safety gaps in large language models, the incident has negligible direct national security implications.

  • Overall national security impact: Minor
  • Response level: Moderate
  • Scope: Multiple nations
  • Primary target: No clear primary
  • Other affected: Unknown

Threat characteristics

  • Imminence: Long-term. The incident represents an ongoing strategic concern regarding model biases and safety guardrails rather than an active national security crisis.
  • Autonomy: Human-supervised. The AI autonomously generated text outputs based on human prompts, with the developer later implementing content filters to supervise and restrict outputs.
  • Novelty: Evolved capability. This represented a significant evolution in the coherence and complexity of toxic outputs generated by large language models.

Impact by dimension

  • Physical security: Negligible. No physical infrastructure, kinetic weapons, or critical systems were targeted or affected in this incident.
  • Information security: Negligible. The incident involved offensive outputs from a writing assistant app, with no indication of coordinated information warfare or intelligence compromise.
  • Sovereignty: Negligible. No sovereign government functions, decision-making processes, or electoral systems were compromised or targeted.
  • Economic security: Negligible. There were no reported threats to strategic industries, financial systems, or critical supply chains.
  • Societal stability: Minor. The AI generated highly offensive, racist content and encouraged suicide, causing psychological distress to users, but lacked the scale to threaten national societal stability.
Explore in the interactive Incident Tracker