Facebook AI Reportedly Mislabels Auschwitz Photos as 'Bullying' and 'Nudity'

Meta's content moderation AI incorrectly flagged and demoted historical photographs from the Auschwitz Memorial Museum, labeling them as 'bullying' and 'nudity'. One image of orphans was deleted by the system. Meta later apologized, stating the notices were sent in error and that the content did not violate policies.

Facebook's AI system allegedly mislabeled around 20 posts from the Auschwitz Memorial Museum as violating community standards for "bullying" and "nudity," reportedly deleting at least one image of orphans. The purported misclassification of historical content allegedly prompted outrage from the museum, which reportedly demanded an explanation. Meta later purportedly apologized, attributing the issue to mistaken notices sent by its AI system and reportedly acknowledged that the posts did not, in fact, violate company policies.

Source: AI Incident Database

Risk classification

  • Primary risk domain: 7 AI system safety, failures, & limitations
  • Primary risk subdomain: 7.3 Lack of capability or robustness

The AI system failed to perform robustly by misclassifying historical memorial photographs as policy violations, demonstrating a lack of capability in contextual understanding.

Additional risk subdomains

  • 1.1 Unfair discrimination and misrepresentation: The algorithm associated historical images of Holocaust victims with offensive categories like bullying and sexual solicitation, misrepresenting the memorial content.

Causal factors

  • Entity: AI
  • Intent: Unintentional
  • Timing: Post-deployment

The incident was caused by an unexpected classification error made by Meta's deployed content moderation AI system.

EU AI Act risk tier

  • Risk tier: 4 Minimal or No Risk

Minimal or No Risk: The system is an automated content moderation filter, which is analogous to a spam filter and poses low direct risk to physical safety, thus falling under Risk Level 4.

AI system and alleged parties

  • AI system: None named
  • AI purpose: Content Moderation; NSFW Content Detection
  • Behaviour type: Autonomous
  • Alleged developer: Meta
  • Alleged deployer: Meta
  • Alleged harmed parties: Survivors of Holocaust victims, General public, Auschwitz Memorial Museum

Harm severity

Highest direct severity in any category: Minor. Severity is scored from Negligible to Catastrophic in each harm category, for harm the reports describe as caused directly or indirectly by the AI system.

  • Physical: direct Negligible, indirect Negligible
  • Infrastructure: direct Negligible, indirect Negligible
  • Property: direct Negligible, indirect Negligible
  • Financial: direct Negligible, indirect Negligible
  • Environmental: direct Negligible, indirect Negligible
  • Malicious content: direct Negligible, indirect Negligible
  • Differential treatment: direct Minor, indirect Negligible
  • Civil rights: direct Negligible, indirect Negligible
  • Democracy: direct Negligible, indirect Negligible
  • Privacy: direct Negligible, indirect Negligible
  • Psychological: direct Minor, indirect Negligible
  • Epistemic: direct Negligible, indirect Negligible
  • Child sexual exploitation and abuse: direct Negligible, indirect Negligible

Differential treatment

Reported: Yes, a representative accused the algorithm of treating genuine Holocaust history with 'suspicion'.

Directly caused: The AI system treated historical posts commemorating Jewish victims of the Holocaust differently by flagging them as policy violations.

Indirectly caused: N/A

Inferred additional harm: N/A

Psychological

Reported: Yes, the report describes the museum staff and community finding the automated censorship offensive and unacceptable.

Directly caused: Museum staff and the broader community experienced distress and offense due to historical victims being labeled as associated with 'bullying' and 'sexual solicitation'.

Indirectly caused: N/A

Inferred additional harm: N/A

People affected

  • Occurrences reported: 1
  • People reportedly exposed: 10

Potential causes

Management

  • Inadequate Risk Assessment: Meta failed to protect historical institutions from automated moderation errors.

Technology

  • Algorithmic Misclassification: Meta's moderation algorithm mistook historical photos for policy violations.
  • Erroneous Notification System: System mistakenly sent violation notices for posts that were not demoted.

Data Inputs

  • Lack of Historical Context: The AI lacked the contextual data to distinguish history from violations.

Human Factors

  • Overreliance on Automated Systems: Lack of human oversight allowed automated flags to be sent to the museum.

Process and Methods

  • Inadequate Content Review: Automated review lacked nuanced evaluation for historical documentation.
  • Lack of Recourse Options: A post was removed summarily without immediate possibility of appeal.

Information quality

  • Classification confidence: High
  • Reason for confidence: The report clearly details the incident where Meta's automated moderation system flagged and deleted historical posts from the Auschwitz Memorial Museum. The roles of the AI, the developer (Meta), and the affected party are explicitly stated, and Meta's apology confirms the system's erroneous actions.
  • Ambiguities identified: There is a slight contradiction where the museum claims 21 posts were flagged and one deleted, while Meta claims the content was never actually demoted in feeds despite notices being sent.

Meta's automated content moderation system mistakenly flagged and deleted historical Holocaust memorial posts from the Auschwitz Memorial Museum's Facebook page. While the incident briefly restricted digital expression and historical preservation, it was an unintentional corporate moderation error with negligible to minor national security implications, and the content was quickly restored.

  • Overall national security impact: Minor
  • Response level: Moderate
  • Scope: Single nation
  • Primary target: Poland
  • Alleged perpetrator: Unknown

Threat characteristics

  • Imminence: Long-term. The incident has been resolved with Meta restoring the content, representing a historical event with no immediate ongoing threat.
  • Autonomy: Full autonomy. The AI system autonomously scanned, flagged, and deleted the content without direct human intervention prior to the enforcement action.
  • Novelty: Established threat. Social media moderation algorithms generating false positives and erroneously flagging historical or sensitive content is an established and frequent issue.

Impact by dimension

  • Physical security: Negligible. No physical systems, critical infrastructure, or kinetic assets were impacted or targeted in this incident.
  • Information security: Minor. The incident involves algorithmic censorship of historical data, but it was an unintentional technical error by a private platform rather than a coordinated state-sponsored information operation.
  • Sovereignty: Negligible. No core government functions, sovereign decision-making processes, or electoral systems were disrupted or targeted.
  • Economic security: Negligible. There were no threats to financial systems, critical supply chains, or strategic technological advantages.
  • Societal stability: Minor. The automated removal of Holocaust memorial content temporarily restricted the museum's historical preservation efforts, representing a minor, localized impact on human rights and digital expression.
Explore in the interactive Incident Tracker