Facebook's Automated Moderation Mistakenly Flagged Landmark's Name as Offensive

Facebook's automated content moderation system incorrectly flagged and removed posts from users mentioning the Plymouth Hoe landmark, misidentifying the word 'Hoe' as a misogynistic slur. This resulted in users receiving warnings, content removals, and temporary bans from the platform. Facebook acknowledged the error and apologized for the impact on residents and visitors.

Facebook's automated system mistakenly labelled posts featuring the seafaring landmark Plymouth Hoe as misogynistic.

Source: AI Incident Database

Risk classification

  • Primary risk domain: 7 AI system safety, failures, & limitations
  • Primary risk subdomain: 7.3 Lack of capability or robustness

The automated moderation system failed to perform robustly under varying linguistic contexts, misinterpreting a geographical landmark name as a slur due to a lack of contextual capability.

Causal factors

  • Entity: AI
  • Intent: Unintentional
  • Timing: Post-deployment

The incident was caused by an automated content moderation AI system acting post-deployment, which unintentionally flagged benign historical terms as offensive slurs.

EU AI Act risk tier

  • Risk tier: 4 Minimal or No Risk

Minimal or No Risk: The system is an automated content moderation/spam filter tool used on a social media platform, which poses low or negligible risk to users and society. According to the definitions, 'AI systems used in entertainment or simple applications like spam filters' fall under Risk Level 4.

AI system and alleged parties

  • AI system: Facebook's automated system (Meta)
  • AI purpose: Content Moderation; Hate Speech Detection
  • Behaviour type: Autonomous
  • Alleged developer: Facebook
  • Alleged deployer: Facebook
  • Alleged harmed parties: Plymouth Hoe residents, Facebook users posting about Plymouth Hoe, Facebook users in Plymouth Hoe

Harm severity

Highest direct severity in any category: Negligible. Severity is scored from Negligible to Catastrophic in each harm category, for harm the reports describe as caused directly or indirectly by the AI system.

  • Physical: direct Negligible, indirect Negligible
  • Infrastructure: direct Negligible, indirect Negligible
  • Property: direct Negligible, indirect Negligible
  • Financial: direct Negligible, indirect Negligible
  • Environmental: direct Negligible, indirect Negligible
  • Malicious content: direct Negligible, indirect Negligible
  • Differential treatment: direct Negligible, indirect Negligible
  • Civil rights: direct Negligible, indirect Negligible
  • Democracy: direct Negligible, indirect Negligible
  • Privacy: direct Negligible, indirect Negligible
  • Psychological: direct Negligible, indirect Negligible
  • Epistemic: direct Negligible, indirect Negligible
  • Child sexual exploitation and abuse: direct Negligible, indirect Negligible

People affected

  • Occurrences reported: 1
  • People reportedly harmed: 4
  • People reportedly exposed: 4

Potential causes

Management

  • Inadequate Risk Assessment: Management did not assess the risk of blocking common geographic terms.

Technology

  • Overly Aggressive Filter: The automated moderation system flagged 'Hoe' without analyzing context.
  • Contextual Blindness: AI failed to recognize 'Plymouth Hoe' as a historical geographical location.

Data Inputs

  • Lack of Localized Training Data: Training dataset did not include UK geographical terms like 'Hoe'.
  • Unrefined Slur Dictionary: The dictionary listed 'hoe' as a slur without homonym disambiguation.

Human Factors

  • User Frustration and Confusion: Users had to bypass filters using spaces or spelling variations.
  • Over-reliance on Automation: System administrators trusted the automated flags without manual verification.

Process and Methods

  • Inadequate Post-Deployment Review: Lack of real-time monitoring to catch widespread false positives quickly.
  • Deficient Appeal Resolution: Muted users faced automated bans without rapid human appeal channels.

Regulatory Environment

  • Pressure to Moderate Slurs: External pressure led to aggressive filtering of potential harassment terms.

Information quality

  • Classification confidence: High
  • Reason for confidence: The report clearly outlines the nature of the AI failure, the specific platform involved, the real-world impact on users, and the developer's response. There is no conflicting information or significant ambiguity regarding what transpired.

An automated content moderation error on Facebook temporarily restricted UK users for mentioning the landmark 'Plymouth Hoe', misidentifying it as a slur. The incident represents a minor commercial platform glitch with negligible national security implications.

  • Overall national security impact: Negligible
  • Response level: Minor
  • Scope: Single nation
  • Primary target: United Kingdom
  • Alleged perpetrator: Unknown

Threat characteristics

  • Imminence: Long-term. The incident represents a minor, resolved software error with no active threat or imminent crisis requiring national response.
  • Autonomy: Full autonomy. The automated content moderation system flagged content and issued temporary bans autonomously without immediate human-in-the-loop verification.
  • Novelty: Established threat. Linguistic false positives and automated moderation errors on social media platforms are well-documented, common occurrences.

Impact by dimension

  • Physical security: Negligible. The incident involved an automated social media content moderation error with no impact on physical systems, critical infrastructure, or human safety.
  • Information security: Negligible. No evidence of coordinated disinformation, intelligence compromise, or foreign influence operations; the event was a localized algorithmic glitch.
  • Sovereignty: Negligible. Core government operations, electoral systems, and state sovereignty were entirely unaffected by this platform-level moderation error.
  • Economic security: Negligible. No strategic technology theft, financial system impacts, or critical supply chain disruptions occurred as a result of this incident.
  • Societal stability: Negligible. While some users experienced temporary account restrictions, the incident was a minor, localized commercial platform error with no systemic impact on civil liberties or societal stability.
Explore in the interactive Incident Tracker