Tiny Changes Let False Claims About COVID-19, Voting Evade Facebook Fact Checks

An analysis by the advocacy group Avaaz found that Facebook's AI-driven misinformation detection system failed to label 42% of posts containing debunked claims about COVID-19 and the 2020 U.S. election. The study revealed that bad actors could easily bypass the system by making minor modifications to images or text, such as changing fonts or backgrounds. These unlabeled posts reached an estimated 142 million views, highlighting significant gaps in the platform's automated moderation capabilities.

Avaaz, an international advocacy group, released a review of Facebook's misinformation identifying software showing that the labeling process failed to label 42% of false information posts, most surrounding COVID-19 and the 2020 USA Presidential Election.

Source: AI Incident Database

Risk classification

  • Primary risk domain: 7 AI system safety, failures, & limitations
  • Primary risk subdomain: 7.3 Lack of capability or robustness

The AI system failed to perform reliably under minor variations in input data, such as font changes or image cropping, allowing debunked misinformation to bypass detection.

Additional risk subdomains

  • 3.1 False or misleading information: The AI's robustness failure directly enabled the widespread dissemination of false and misleading information regarding COVID-19 and the U.S. election.

Causal factors

  • Entity: AI
  • Intent: Unintentional
  • Timing: Post-deployment

The incident was caused by the failure of Facebook's deployed AI duplicate detection system to robustly identify and label modified misinformation posts.

EU AI Act risk tier

  • Risk tier: 4 Minimal or No Risk

Minimal or No Risk: The AI system is an automated content moderation filter, which is analogous to 'simple applications like spam filters' and poses minimal direct risk under the EU AI Act framework.

AI system and alleged parties

  • AI system: Facebook misinformation identifying software (Meta)
  • AI purpose: Content Moderation; Content Verification
  • Behaviour type: Tool
  • Alleged developer: Facebook
  • Alleged deployer: Facebook
  • Alleged harmed parties: Facebook users interested in the US Presidential Election, Facebook users interested in COVID information, Facebook users

Harm severity

Highest direct severity in any category: Severe. Severity is scored from Negligible to Catastrophic in each harm category, for harm the reports describe as caused directly or indirectly by the AI system.

  • Physical: direct Negligible, indirect Negligible
  • Infrastructure: direct Negligible, indirect Negligible
  • Property: direct Negligible, indirect Negligible
  • Financial: direct Negligible, indirect Negligible
  • Environmental: direct Negligible, indirect Negligible
  • Malicious content: direct Severe, indirect Negligible
  • Differential treatment: direct Negligible, indirect Negligible
  • Civil rights: direct Negligible, indirect Negligible
  • Democracy: direct Minor, indirect Negligible
  • Privacy: direct Negligible, indirect Negligible
  • Psychological: direct Negligible, indirect Negligible
  • Epistemic: direct Severe, indirect Negligible
  • Child sexual exploitation and abuse: direct Negligible, indirect Negligible

Malicious content

Reported: The report explicitly describes the spread of malicious misinformation and hoaxes.

Directly caused: The AI system's failure directly allowed the spread of 738 unlabeled posts containing debunked claims, which received an estimated 142 million views.

Indirectly caused: N/A

Inferred additional harm: It is likely that many more modified posts bypassed the system beyond the sample analyzed by Avaaz, exposing millions more users to toxic or malicious misinformation.

Democracy

Reported: The report explicitly describes the spread of misinformation aimed at undermining democratic norms, specifically regarding the U.S. election.

Directly caused: The failure to label election-related misinformation directly allowed false claims about voting procedures, such as mail-in ballots needing two stamps, to circulate without warning.

Indirectly caused: N/A

Inferred additional harm: The widespread circulation of unlabeled election misinformation likely eroded trust in democratic processes and voter participation among exposed users.

Epistemic

Reported: The report explicitly describes epistemic harm through the spread of debunked claims and hoaxes about COVID-19 and the U.S. election.

Directly caused: The AI's failure to label 42% of debunked posts directly resulted in 142 million views of false information without warning labels.

Indirectly caused: N/A

Inferred additional harm: The failure likely contributed to a broader erosion of shared truth and increased belief in conspiracy theories regarding public health and elections.

People affected

  • Occurrences reported: 1
  • People reportedly exposed: 142000000

Potential causes

Management

  • Failure to Mitigate Known Gaps: Management failed to close loopholes despite knowing actors bypass systems.
  • Inadequate Policy Enforcement: Enforcement actions against repeat offender pages were inconsistent.

Technology

  • AI Sensitivity to Minor Tweaks: AI fails to match near-identical images with minor crops or background changes.
  • Text and Font Recognition Limits: Altering fonts or converting image text to status updates bypasses detection.
  • Inability to Track Mutated Memes: System cannot identify mutated versions of previously debunked viral memes.

Data Inputs

  • Incomplete Reference Database: Lack of diverse variations of debunked content in the match database.
  • Unlabeled Variation Data: Slightly altered posts lack the labels needed to train detection models.
  • Delayed Fact-Checking Inputs: Fact-checker data arrives after variations have already gone viral.

Human Factors

  • Orchestrated Adversarial Behavior: Bad actors actively edit content to deliberately exploit system vulnerabilities.
  • User Sharing Behaviors: Users rapidly share unlabeled misinformation, increasing viral reach.
  • Manual Review Bottlenecks: Human reviewers cannot keep pace with the volume of mutated variations.

Process and Methods

  • Inconsistent Labeling Workflows: Identical claims receive different warning labels or none at all.
  • Flawed Duplication Matching: The process for matching duplicate posts fails when minor edits are made.
  • Inadequate Repeat Offender Tracking: Failure to consistently flag pages that repeatedly post misinformation.

Regulatory Environment

  • Lack of External Standards: Absence of regulatory standards for social media misinformation detection.

Information quality

  • Classification confidence: High
  • Reason for confidence: The report clearly outlines the AI system's role, the specific failure mode (lack of robustness to minor image/text modifications), and provides quantitative data from a structured study by Avaaz.
  • Ambiguities identified: The exact technical architecture of Facebook's AI duplicate detection system is not detailed, and 'views' is used as a proxy for user exposure.
  • Alternative interpretations: The incident could be viewed purely as a human moderation policy failure rather than an AI failure, but the report explicitly highlights the AI's inability to detect modified duplicates as the core technical gap.

The failure of Facebook's AI-driven moderation system to detect modified duplicates of debunked misinformation allowed false claims about COVID-19 and the 2020 U.S. election to reach an estimated 142 million views. This vulnerability highlights how simple adversarial perturbations can bypass automated defenses, presenting notable challenges to information security, societal stability, and democratic processes.

  • Overall national security impact: Substantial
  • Response level: Substantial
  • Scope: Multiple nations
  • Primary target: United States
  • Other affected: Global Facebook users
  • Alleged perpetrator: Unknown

Threat characteristics

  • Imminence: Long-term. Represents an ongoing strategic concern regarding platform vulnerability to adversarial manipulation rather than an active, immediate crisis requiring 72-hour intervention.
  • Autonomy: Human-supervised. The moderation system operates autonomously to identify and label posts, but operates under a broader framework of human oversight and manual fact-checking partnerships.
  • Novelty: Evolved capability. While bypassing content filters is an established challenge, the systematic exploitation of simple image and text modifications at this scale represents an evolved threat capability.

Impact by dimension

  • Physical security: Negligible. No direct threats to physical systems, critical infrastructure, or direct human casualties are reported in the incident details.
  • Information security: Substantial. Bad actors exploited vulnerabilities in automated moderation to spread election and public health misinformation, reaching 142 million views, representing a substantial gap in defending against information operations.
  • Sovereignty: Substantial. The failure allowed false claims regarding voting procedures, such as mail-in ballots requiring extra postage, to circulate widely during the 2020 U.S. presidential election, impacting democratic processes.
  • Economic security: Negligible. The incident does not involve threats to financial systems, strategic technology theft, or loss of technological competitive advantage.
  • Societal stability: Substantial. Large-scale spread of debunked COVID-19 and election claims eroded public trust and societal cohesion, exploiting algorithmic vulnerabilities to bypass safety guardrails.
Explore in the interactive Incident Tracker