AI deepfake detection tools are exhibiting significant performance biases, as they are primarily trained on Western-centric data. This results in high error rates when analyzing content from the Global South, including false positives for non-native English speakers and failures to recognize lower-quality media common in those regions. These inaccuracies undermine election integrity and hinder the ability of journalists and researchers to combat misinformation effectively.
AI deepfake detection tools are reportedly failing voters in the Global South due to biases in their training data. These tools, which prioritize English language and Western faces, show reduced accuracy when detecting manipulated content from non-Western regions. As a result of this detection gap, election integrity faces threats from and the amplification of misinformation, which leaves journalists and researchers with inadequate resources to combat the issue.
Risk classification
- Primary risk domain: 1 Discrimination & Toxicity
- Primary risk subdomain: 1.3 Unequal performance across groups
The AI detection tools perform with significantly lower accuracy and higher false-positive rates for non-Western faces and non-native English speakers due to biased training data.
Additional risk subdomains
- 7.3 Lack of capability or robustness: The detection models fail to perform robustly under real-world conditions such as low-quality media, background noise, and compression.
- 4.1 Disinformation, surveillance, and influence at scale: The failure of detection tools in the Global South allows political disinformation campaigns to spread unchecked during elections.
Causal factors
- Entity: AI
- Intent: Unintentional
- Timing: Post-deployment
The incident is caused by biased classification decisions made by deployed AI detection systems, which unintentionally fail to perform robustly on non-Western data.
EU AI Act risk tier
- Risk tier: 3 Limited Risk
Limited Risk: The report describes AI systems used for deepfake detection and synthetic media verification, which fall under the category of managing AI-generated content and deepfakes.
AI system and alleged parties
- AI system: Reality Defender, True Media's detection tool, unspecified AI deepfake detection tools (Reality Defender, True Media)
- AI purpose: Content Verification; Image Verification
- Behaviour type: Tool
- Alleged developer: Unknown deepfake detection technology developers, True Media, Reality Defender
- Alleged deployer: Unknown deepfake detection technology developers, True Media, Reality Defender
- Alleged harmed parties: Political researchers, Non-native English speakers, Global South local fact-checkers, Global South journalists, Global South Citizens, Civil society organizations in developing countries
Harm severity
Highest direct severity in any category: Severe. Severity is scored from Negligible to Catastrophic in each harm category, for harm the reports describe as caused directly or indirectly by the AI system.
- Physical: direct Negligible, indirect Negligible
- Infrastructure: direct Negligible, indirect Negligible
- Property: direct Negligible, indirect Negligible
- Financial: direct Negligible, indirect Negligible
- Environmental: direct Negligible, indirect Negligible
- Malicious content: direct Negligible, indirect Minor
- Differential treatment: direct Minor, indirect Minor
- Civil rights: direct Negligible, indirect Minor
- Democracy: direct Negligible, indirect Minor
- Privacy: direct Negligible, indirect Negligible
- Psychological: direct Negligible, indirect Negligible
- Epistemic: direct Minor, indirect Substantial
- Child sexual exploitation and abuse: direct Negligible, indirect Negligible
Malicious content
Reported: Yes, the report describes the spread of AI-generated political disinformation and deepfakes, such as manipulated images of Taylor Swift fans supporting Donald Trump.
Directly caused: N/A
Indirectly caused: The failure of AI detection tools indirectly facilitates the spread of synthetic media and political disinformation across platforms in the Global South.
Inferred additional harm: It is highly likely that millions of social media users in the Global South are exposed to undetected toxic, manipulated, or malicious political content due to these detection failures.
Differential treatment
Reported: Yes, the report explicitly describes how AI detection tools perform significantly worse for non-Western faces, non-native English speakers, and lower-quality media common in the Global South.
Directly caused: The AI models directly exhibit performance bias, returning high rates of false positives for non-native English speakers and failing to accurately detect deepfakes of non-white subjects.
Indirectly caused: Journalists and researchers in the Global South are left with fewer resources and inaccurate tools compared to their Western counterparts, hindering their ability to combat disinformation.
Inferred additional harm: This systemic bias likely leads to unequal protection from disinformation for billions of people living outside of Western markets.
Civil rights
Reported: Yes, the report mentions that Witness helps people use technology to support human rights, and notes that faulty models could encourage legislators to crack down on imaginary problems, potentially restricting civil rights.
Directly caused: N/A
Indirectly caused: The inability to detect deepfakes hinders the protection of human rights, and false positives could lead to regulatory crackdowns that restrict freedom of expression.
Inferred additional harm: The spread of undetected political deepfakes likely interferes with citizens' civil rights to participate in free and fair elections.
Democracy
Reported: Yes, the report describes how the use of generative AI for political purposes and the difficulty of detecting it leaves journalists and researchers unable to address the deluge of disinformation heading their way during elections.
Directly caused: N/A
Indirectly caused: The failure of detection tools allows political disinformation to spread during elections, undermining democratic processes and public trust in information ecosystems.
Inferred additional harm: The integrity of democratic elections in multiple countries in the Global South is likely compromised by the unchecked spread of synthetic political media.
Epistemic
Reported: Yes, the report explicitly describes how the failure of detection tools leaves journalists and researchers unable to address the deluge of disinformation, leading to false positives and false negatives.
Directly caused: The biased tools directly generate epistemic confusion by falsely flagging authentic non-native English text as AI-generated (false positives) and failing to identify actual deepfakes (false negatives).
Indirectly caused: This undermines the overall information ecosystem, making it difficult for the public to distinguish between real and synthetic media, thereby eroding shared truth.
Inferred additional harm: Widespread epistemic confusion and a loss of public trust in news media are highly likely across affected regions in the Global South.
People affected
- Occurrences reported: 1
- People reportedly harmed: 100
- People reportedly exposed: 100
Potential causes
Management
- Misallocated safety funding: Funding goes to detection instead of resilient news ecosystems.
- Prohibitive costs of commercial tools: High cost of reliable tools forces reliance on inaccurate free versions.
Technology
- Lack of local compute infrastructure: Global South lacks data centers to run local detection models.
- Sensitivity to media compression: Compression and background noise trigger false detection results.
- Western-biased model design: Detection tools are optimized for Western markets and demographics.
Data Inputs
- Undigitised local data sources: African data in hard copy prevents training models on local contexts.
- Lack of diverse training data: Models lack training on non-Western accents, languages, and faces.
- Training on high-quality media: Models fail on low-quality photos from cheap smartphones.
Human Factors
- Untrained researchers: Researchers mistake simple cheapfakes for AI manipulations.
- Reliance on inaccurate free tools: Journalists use faulty free tools due to lack of better options.
Process and Methods
- Significant verification lag time: Sending media overseas takes weeks, allowing disinformation to spread.
- Overwhelmed rapid response teams: High volume of cases exceeds the capacity of verification programs.
Regulatory Environment
- Risk of misguided regulation: Inflated false positives risk driving legislation for imaginary issues.
Information quality
- Classification confidence: High
- Reason for confidence: The report provides clear, detailed, and consistent testimonies from multiple experts and organizations (Witness, Tech Global Institute, Thraets) regarding the specific technical biases and real-world impacts of AI detection tools in the Global South. The causal mechanisms (biased training data, lack of local compute, low-quality hardware inputs) are explicitly explained.
- Ambiguities identified: The report does not provide quantitative data on the exact error rates or the precise number of false positives/negatives occurring in specific elections.
- Alternative interpretations: The failures could be interpreted primarily as a market access or economic inequality issue rather than an inherent AI safety failure, though the resulting algorithmic bias directly constitutes an AI safety risk.
Technical bias in AI deepfake detection tools, which are trained primarily on Western data, leads to severe performance degradation in the Global South. This failure undermines the ability of researchers and journalists to counter political disinformation during elections, posing a substantial threat to information security and democratic processes in affected nations.
- Overall national security impact: Substantial
- Response level: Substantial
- Scope: Multiple nations
- Primary target: No clear primary
- Other affected: Global South nations including Bangladesh, Senegal, and Ghana
- Alleged perpetrator: Unknown
Threat characteristics
- Imminence: Long-term. This represents an ongoing strategic concern regarding systemic algorithmic bias and global technological inequality rather than an immediate crisis.
- Autonomy: Human-controlled. The AI tools function as classifiers to assist human journalists and researchers, who ultimately make the final verification decisions.
- Novelty: Established threat. Algorithmic bias due to unrepresentative training data is a well-documented and established threat in AI systems.
Impact by dimension
- Physical security: Negligible. No physical threat to systems, infrastructure, or human safety is indicated in the incident report.
- Information security: Substantial. Western-centric biases in deepfake detection tools hinder the identification of synthetic media, allowing political disinformation campaigns to spread unchecked in non-Western regions.
- Sovereignty: Substantial. The failure of verification tools during elections undermines democratic processes and public trust in electoral systems across multiple sovereign nations.
- Economic security: Minor. Local organizations face economic barriers due to the prohibitive costs of proprietary Western verification tools, though strategic economic systems remain unaffected.
- Societal stability: Substantial. Systemic performance bias results in unequal protection from disinformation, potentially leading to social manipulation at scale and regulatory overreach that restricts civil liberties.