In March 2020, Facebook's automated anti-spam system experienced a bug that caused it to incorrectly flag and remove legitimate news articles about the COVID-19 pandemic. Users reported being unable to share links from various news outlets, with the platform citing violations of community standards on spam. Facebook acknowledged the issue as a system bug and worked to restore the affected posts.
Facebook was reported by users for blocking posts of legitimate news about the coronavirus pandemic, allegedly due to a bug in an anti-spam system.
Risk classification
- Primary risk domain: 7 AI system safety, failures, & limitations
- Primary risk subdomain: 7.3 Lack of capability or robustness
The incident represents a failure of the AI system to perform reliably under varying conditions, resulting in false-positive spam classifications for legitimate news articles.
Causal factors
- Entity: AI
- Intent: Unintentional
- Timing: Post-deployment
The incident was caused by an unexpected bug in Facebook's deployed machine learning anti-spam system, which unintentionally blocked legitimate news articles.
EU AI Act risk tier
- Risk tier: 4 Minimal or No Risk
Minimal or No Risk: The system is described as an 'anti-spam system' and 'spam filters', which are explicitly categorized as posing low or negligible risks under the EU AI Act.
AI system and alleged parties
- AI system: anti-spam system (Meta)
- AI purpose: Spam Filtering; Content Moderation
- Behaviour type: Autonomous
- Alleged developer: Facebook
- Alleged deployer: Facebook
- Alleged harmed parties: Facebook users posting legitimate COVID-19 news, Facebook users
Harm severity
Highest direct severity in any category: Minor. Severity is scored from Negligible to Catastrophic in each harm category, for harm the reports describe as caused directly or indirectly by the AI system.
- Physical: direct Negligible, indirect Negligible
- Infrastructure: direct Negligible, indirect Negligible
- Property: direct Negligible, indirect Negligible
- Financial: direct Negligible, indirect Negligible
- Environmental: direct Negligible, indirect Negligible
- Malicious content: direct Negligible, indirect Negligible
- Differential treatment: direct Negligible, indirect Negligible
- Civil rights: direct Negligible, indirect Negligible
- Democracy: direct Negligible, indirect Negligible
- Privacy: direct Negligible, indirect Negligible
- Psychological: direct Negligible, indirect Negligible
- Epistemic: direct Negligible, indirect Negligible
- Child sexual exploitation and abuse: direct Negligible, indirect Negligible
People affected
- Occurrences reported: 1
- People reportedly harmed: 3
- People reportedly exposed: 3
Potential causes
Management
- Moderator Workforce Reduction: Management sent moderators home due to the COVID-19 pandemic.
- Inadequate Contingency Planning: Lack of backup plans for sudden loss of human content moderators.
Technology
- Anti-Spam Algorithmic Bug: An anti-spam rule went haywire, misclassifying news as spam.
- Machine Learning Over-reliance: System relied on automated ML which lacked nuance without human oversight.
Data Inputs
- Misclassification of News Links: Legitimate news URLs were incorrectly flagged as spam signatures.
- Lack of Contextual Data: The spam filter could not differentiate spam from high-quality COVID news.
Human Factors
- Lack of Real-Time Human Oversight: Fewer content moderators were active to override incorrect automated flags.
- Inability to Work from Home: Moderators could not work remotely due to privacy commitments.
Process and Methods
- Inflexible Content Review Rules: Automated anti-spam rules lacked the flexibility to adapt to pandemic news.
- Sudden Transition to Automation: Rapid shift to automated moderation lacked adequate transition protocols.
Regulatory Environment
- Strict Privacy Commitments: Privacy commitments prevented moderators from working from home.
Information quality
- Classification confidence: High
- Reason for confidence: The reports clearly identify the AI system, the nature of the failure, and the immediate impact. Facebook executives confirmed the bug, reducing ambiguity about whether it was an intentional policy change or a system error.
- Ambiguities identified: The exact technical cause of the bug is not specified, and there is a minor contradiction between a former executive's speculation about WFH content moderators and Facebook's official denial.
- Alternative interpretations: None that are highly plausible, as both users and Facebook confirmed it was an automated system error.
An unintentional bug in Facebook's automated anti-spam system temporarily blocked users from sharing legitimate COVID-19 news articles during the onset of the pandemic. While causing temporary public confusion and restricting access to critical information, the incident was quickly resolved and has negligible long-term national security implications.
- Overall national security impact: Minor
- Response level: Moderate
- Scope: Multiple nations
- Primary target: No clear primary
- Other affected: Global
- Alleged perpetrator: Unknown
Threat characteristics
- Imminence: Long-term. The incident was a temporary system bug that was quickly resolved, representing an ongoing strategic concern regarding automated content moderation.
- Autonomy: Full autonomy. The anti-spam system acted autonomously to flag and block content in real-time without direct human oversight.
- Novelty: Established threat. False positives in automated content moderation systems are a well-documented and established technical issue.
Impact by dimension
- Physical security: Negligible. No physical systems, critical infrastructure, or kinetic assets were compromised or targeted in this incident.
- Information security: Minor. Unintentional system bug temporarily blocked legitimate COVID-19 news, but was not an adversarial information warfare campaign or intelligence breach.
- Sovereignty: Negligible. Private platform content moderation bug with no direct threat to state authority, electoral systems, or core government operations.
- Economic security: Negligible. No strategic technology theft, financial system attacks, or significant economic security impacts occurred.
- Societal stability: Minor. Temporarily restricted access to critical public health information during a pandemic, potentially causing public confusion, but resolved quickly.