During the May 2021 Israeli-Palestinian conflict, major social media platforms including Facebook, Instagram, and Twitter experienced widespread failures in their automated content moderation systems. These AI-driven systems incorrectly flagged and removed millions of posts and accounts, particularly those documenting the conflict from a Palestinian perspective, often misidentifying protest imagery as harassment or political hashtags as terrorist-affiliated. Additionally, the payment app Venmo mistakenly suspended humanitarian aid transactions. The companies attributed these issues to algorithmic glitches and over-enforcement, leading to significant criticism regarding the impact of automated moderation on free speech and marginalized groups.
Facebook, Instagram, and Twitter wrongly blocked or restricted millions of pro-Palestinian posts and accounts related to the Israeli-Palestinian conflict, citing errors in their automated content moderation system.
Risk classification
- Primary risk domain: 1 Discrimination & Toxicity
- Primary risk subdomain: 1.3 Unequal performance across groups
The content moderation algorithms exhibited unequal performance, disproportionately flagging and censoring Palestinian voices, Arabic terms, and protest documentation compared to other groups.
Additional risk subdomains
- 7.3 Lack of capability or robustness: The automated moderation systems lacked the capability and robustness to accurately distinguish between spam/terrorism and legitimate political expression at scale.
Causal factors
- Entity: AI
- Intent: Unintentional
- Timing: Post-deployment
The incident was caused by deployed content moderation and transaction monitoring AI systems that unintentionally over-enforced policies, leading to unexpected blocking of legitimate content and transactions.
EU AI Act risk tier
- Risk tier: 4 Minimal or No Risk
Minimal or No Risk: The AI systems described are automated content moderation and spam filtering systems, which are classified as posing minimal risk under the EU AI Act. The report notes that 'Twitter said its service mistakenly identified the rapid-firing tweeting... as spam' and refers to the incident as a 'spam filter issue'.
AI system and alleged parties
- AI system: automated content moderation system
- AI purpose: Content Moderation; Spam Filtering
- Behaviour type: Autonomous
- Alleged developer: Twitter, Instagram, Facebook
- Alleged deployer: Twitter, Instagram, Facebook
- Alleged harmed parties: Twitter Users, Palestinian social media users, Instagram users, Facebook users, Facebook employees having families affected by the conflict
Harm severity
Highest direct severity in any category: Severe. Severity is scored from Negligible to Catastrophic in each harm category, for harm the reports describe as caused directly or indirectly by the AI system.
- Physical: direct Negligible, indirect Negligible
- Infrastructure: direct Negligible, indirect Negligible
- Property: direct Negligible, indirect Negligible
- Financial: direct Minor, indirect Negligible
- Environmental: direct Negligible, indirect Negligible
- Malicious content: direct Negligible, indirect Minor
- Differential treatment: direct Substantial, indirect Negligible
- Civil rights: direct Severe, indirect Substantial
- Democracy: direct Negligible, indirect Negligible
- Privacy: direct Negligible, indirect Negligible
- Psychological: direct Negligible, indirect Negligible
- Epistemic: direct Negligible, indirect Minor
- Child sexual exploitation and abuse: direct Negligible, indirect Negligible
Financial
Reported: Yes, the report notes that Venmo mistakenly suspended transactions of humanitarian aid to Palestinians during the war.
Directly caused: The transaction monitoring AI directly suspended humanitarian aid transactions, delaying or blocking financial support to those in need, though the exact monetary value is not specified in the report.
Indirectly caused: N/A
Inferred additional harm: The suspension of aid transactions likely caused indirect financial hardship and delayed critical resources for humanitarian organizations and recipients in the conflict zone.
Malicious content
Reported: Yes, the report notes that the automated systems sometimes under-enforced, 'allowing harmful misinformation and violent and hateful language to proliferate'.
Directly caused: N/A
Indirectly caused: The failure of the moderation algorithms to properly balance enforcement led to the proliferation of violent and hateful language on the platforms during the conflict.
Inferred additional harm: The proliferation of unmoderated toxic content likely reached millions of users, potentially exacerbating real-world tensions and hostility.
Differential treatment
Reported: Yes, the report explicitly describes how the automated systems wrongly penalized marginalized groups, specifically Palestinians and Black Americans.
Directly caused: The AI systems disproportionately blocked and restricted pro-Palestinian accounts, protest images, and Arabic cultural terms, while failing to apply the same level of restriction to other groups.
Indirectly caused: N/A
Inferred additional harm: The systemic bias in content moderation algorithms likely results in ongoing differential treatment and silencing of marginalized groups globally during crises.
Civil rights
Reported: Yes, the report describes the suppression of free speech and digital rights for millions of users.
Directly caused: The automated systems directly violated users' freedom of expression by wrongly removing millions of posts and restricting accounts documenting a global crisis.
Indirectly caused: The censorship hindered the ability of activists and journalists to document potential human rights abuses and share critical information during the conflict.
Inferred additional harm: The widespread suppression of digital expression during a crisis likely had a chilling effect on free speech and human rights advocacy across the region.
Epistemic
Reported: Yes, the report mentions that the systems under-enforced, allowing 'harmful misinformation' to proliferate.
Directly caused: N/A
Indirectly caused: The failure of the moderation algorithms allowed false or misleading information about the conflict and other topics to spread, corrupting the information environment.
Inferred additional harm: The spread of unmoderated misinformation likely misled public opinion and distorted the shared understanding of the conflict.
People affected
- Occurrences reported: 1
- People reportedly harmed: 2000000
- People reportedly exposed: 2000000
Potential causes
Management
- Prioritizing Automation: Companies shifted from human moderators to policing algorithms over time.
- Lack of Transparency: Management failed to disclose when governments secretly referred accounts.
Technology
- Spam Detection Overcorrection: Twitter AI mistook rapid-fire tweeting during the conflict as spam activity.
- Hate Speech Algorithm Glitch: Facebook AI misidentified the Al-Aqsa Mosque hashtag as a terrorist group.
- Image Classifier Mislabeling: AI mistakenly labeled images of protests as harassment or bullying.
Data Inputs
- Keyword Association Error: The name Qassam triggered blocks due to association with Hamas military wing.
- Contextual Keyword Triggers: Keywords like Zionist are flagged as hate speech without proper context.
- Lack of Context in Language Data: Algorithms failed to differentiate cultural discourse from terrorist terms.
Human Factors
- Human Classification Error: Facebook employees mistakenly designated Al-Aqsa Mosque as terrorist group.
- Over-reliance on Automation: Platforms relied on policing algorithms instead of human decision-makers.
- Lack of Regional Representation: Review teams lacked Palestinian representation to provide proper context.
Process and Methods
- Inadequate Content Scaling: Algorithmic content moderation failed to work effectively at global scale.
- Government Reporting Channels: Israeli Cyber Unit used direct channels to report and block content.
- Lack of Appeal Recourse: Marginalized groups had little recourse when posts were incorrectly flagged.
Regulatory Environment
- Sanctions Compliance Pressures: Venmo suspended humanitarian aid transactions to comply with US sanctions.
- Government Censorship Requests: Tech companies complied with local laws and government takedown requests.
Information quality
- Classification confidence: High
- Reason for confidence: The reports provide clear, detailed accounts from multiple platforms (Facebook, Instagram, Twitter, Venmo) regarding the failure of their automated moderation and transaction systems. The role of the AI is explicitly stated, and the companies themselves acknowledged the algorithmic errors. While some specific technical details of the algorithms are proprietary, the overall nature of the failure and its impacts are well-documented.
- Ambiguities identified: The exact technical mechanisms of the content moderation algorithms and the specific criteria used by Venmo's transaction monitoring system are not fully detailed.
- Alternative interpretations: Some might argue the content moderation decisions were politically motivated rather than purely algorithmic glitches, though the platforms officially attributed them to AI errors.
During the May 2021 conflict, automated content moderation systems on major platforms failed at scale, suppressing millions of pro-Palestinian posts and disrupting humanitarian aid transactions. While unintentional, the incident highlights how AI-driven information control can severely impact human rights, public communication, and the information environment during active geopolitical crises.
- Overall national security impact: Substantial
- Response level: Substantial
- Scope: Multiple nations
- Primary target: Palestine
- Other affected: Israel, Colombia, Canada, United States
- Alleged perpetrator: Unknown
Threat characteristics
- Imminence: Immediate. The algorithmic suppression occurred in real-time during an active kinetic conflict, rapidly impacting communication and aid distribution.
- Autonomy: Human-supervised. The AI systems automatically flagged and removed content without immediate human intervention, though human operators ultimately intervened to correct the errors.
- Novelty: Established threat. Algorithmic moderation errors and over-enforcement are known issues, though this incident demonstrated their severe impact when scaled during a geopolitical crisis.
Impact by dimension
- Physical security: Negligible. The incident involved content moderation and transaction filtering failures on social media and payment apps, with no physical kinetic impacts or critical infrastructure disruption.
- Information security: Substantial. Automated systems suppressed millions of posts documenting an active conflict and allowed harmful misinformation to spread, heavily disrupting the information environment during a geopolitical crisis.
- Sovereignty: Negligible. No core government functions, sovereign decision-making, or state authority processes were directly compromised or manipulated by these algorithmic failures.
- Economic security: Minor. Venmo mistakenly suspended humanitarian aid transactions, causing localized financial disruptions for recipients, but without broader strategic economic or technological security threats.
- Societal stability: Substantial. Algorithmic over-enforcement systematically suppressed free speech and protest documentation for millions of users, disproportionately affecting Palestinian voices and impacting digital human rights.