Leading AI Models Reportedly Found to Mimic Russian Disinformation in 33% of Cases and to Cite Fake Moscow News Sites

A NewsGuard audit of 10 leading AI chatbots found that they repeated Russian disinformation narratives in approximately 32% of responses. The models frequently cited fake local news sites as reliable sources, authoritatively presenting false claims as fact. This behavior demonstrates how AI can inadvertently validate and amplify disinformation campaigns, posing risks to public information integrity during election cycles.

An audit by NewsGuard revealed that leading chatbots, including ChatGPT-4, You.com’s Smart Assistant, and others, repeated Russian disinformation narratives in one-third of their responses. These narratives originated from a network of fake news sites created by John Mark Dougan (Incident 701). The audit tested 570 prompts across 10 AI chatbots, showing that AI remains a tool for spreading disinformation despite efforts to prevent misuse.

Source: AI Incident Database

Risk classification

  • Primary risk domain: 3 Misinformation
  • Primary risk subdomain: 3.1 False or misleading information

The chatbots repeated and validated false Russian propaganda narratives, presenting fabricated stories and fake news sources as factual information to users.

Additional risk subdomains

  • 4.1 Disinformation, surveillance, and influence at scale: The chatbots amplified coordinated foreign influence campaigns designed to manipulate public opinion and political processes during election cycles.
  • 7.3 Lack of capability or robustness: The AI models failed to perform reliably by failing to recognize propaganda websites and repeating known falsehoods.

Causal factors

  • Entity: AI
  • Intent: Unintentional
  • Timing: Post-deployment

The incident was caused by deployed AI chatbots unintentionally generating and repeating Russian disinformation narratives when prompted by users.

EU AI Act risk tier

  • Risk tier: 3 Limited Risk

Limited Risk: The report describes chatbots and generative AI systems, which are classified as limited risk under the EU AI Act and are subject to transparency obligations to ensure users know they are interacting with an AI.

AI system and alleged parties

  • AI system: ChatGPT-4, Claude, Copilot, Gemini, Grok, Meta AI, Perplexity's answer engine, Pi, You.com’s Smart Assistant, le Chat (Anthropic, Google, Inflection, Meta, Microsoft, Mistral AI, OpenAI, Perplexity AI, You.com, xAI)
  • AI purpose: Chatbot; Question Answering
  • Behaviour type: Assistant
  • Alleged developer: You.com, xAI, Perplexity, OpenAI, Mistral, Microsoft, Meta, Inflection, Google, Anthropic
  • Alleged deployer: You.com, xAI, Perplexity, OpenAI, Mistral, Microsoft, Meta, John Mark Dougan, Inflection, Google, Anthropic
  • Alleged harmed parties: Western democracies, Volodymyr Zelenskyy, Ukraine, Secret Service, Researchers, Media consumers, General public, Electoral integrity, AI companies facing reputational damage

Harm severity

Highest direct severity in any category: Severe. Severity is scored from Negligible to Catastrophic in each harm category, for harm the reports describe as caused directly or indirectly by the AI system.

  • Physical: direct Negligible, indirect Negligible
  • Infrastructure: direct Negligible, indirect Negligible
  • Property: direct Negligible, indirect Negligible
  • Financial: direct Negligible, indirect Negligible
  • Environmental: direct Negligible, indirect Negligible
  • Malicious content: direct Minor, indirect Negligible
  • Differential treatment: direct Negligible, indirect Negligible
  • Civil rights: direct Negligible, indirect Negligible
  • Democracy: direct Minor, indirect Minor
  • Privacy: direct Negligible, indirect Negligible
  • Psychological: direct Negligible, indirect Negligible
  • Epistemic: direct Minor, indirect Substantial
  • Child sexual exploitation and abuse: direct Negligible, indirect Negligible

Malicious content

Reported: Yes, the report describes the generation and spread of Russian disinformation narratives by the chatbots.

Directly caused: The 10 chatbots generated 152 responses containing explicit disinformation and 29 responses repeating false claims with a disclaimer out of 570 total test prompts.

Indirectly caused: N/A

Inferred additional harm: It is likely that similar disinformation is generated and spread to numerous public users querying these chatbots about current events.

Democracy

Reported: Yes, the report highlights concerns about the chatbots spreading election-related misinformation during a global election year.

Directly caused: The chatbots repeated false narratives about U.S. elections, such as a nonexistent Ukrainian troll factory interfering with U.S. elections.

Indirectly caused: The propagation of such narratives could influence public opinion and undermine trust in democratic processes.

Inferred additional harm: Widespread exposure to these narratives could potentially sway voter behavior or erode trust in democratic institutions, though the exact scale is unquantified.

Epistemic

Reported: Yes, the report explicitly details how chatbots present false reports, satire, and fiction as fact, citing fake local news sites as reliable sources.

Directly caused: Chatbots repeated false claims about a wiretap at Mar-a-Lago and the murder of an Egyptian journalist, citing propaganda sites like 'The Boston Times' and 'The Houston Post' as credible.

Indirectly caused: This creates an unvirtuous cycle where falsehoods are generated, repeated, and validated by AI platforms, eroding the shared information ecosystem.

Inferred additional harm: Millions of users relying on these chatbots for news may develop inaccurate beliefs about critical geopolitical and domestic events.

People affected

  • Occurrences reported: 1

Potential causes

Management

  • Prioritizing Growth over Safety: AI companies deployed models without sufficient safety guardrails for news.
  • Inadequate Risk Assessment: Failure to assess risks of Russian disinformation networks on election safety.
  • Unresponsive Safety Teams: AI companies failed to respond to NewsGuard's inquiries regarding findings.

Technology

  • Inability to Detect AI Personas: Chatbots failed to recognize AI-generated personas like Olesya Movchan.
  • Vulnerability to Malign Prompts: Models generated fake whistleblower scripts when prompted by bad actors.
  • Lack of Source Credibility Logic: Chatbots treated fake local news sites as highly credible sources.

Data Inputs

  • Ingestion of Propaganda Outlets: Models trained on or retrieved data from 167 fake Russian news sites.
  • Unverified YouTube Testimonies: Chatbots used fake YouTube whistleblower videos as authoritative sources.
  • Mimicked Trustworthy Domains: Fake sites used names of defunct real newspapers to dupe training filters.

Human Factors

  • Exploitation by Malicious Actors: Malign actors intentionally prompt chatbots to produce convincing fake news.
  • User Overreliance on Chatbots: Users increasingly rely on chatbots for quick and customized news updates.

Process and Methods

  • Inadequate Safeguards for News: AI companies failed to implement special filters for news and election topics.
  • Ineffective Source Verification: Lack of processes to verify the authenticity of cited web domains.
  • Weak Guardrails for Election Info: Pledges to curb election misinformation were not backed by robust actions.

Regulatory Environment

  • Absence of Clear AI Regulations: Governments are still striving to regulate generative AI safety effectively.
  • Political Backlash on Watchdogs: Investigations into watchdogs like NewsGuard complicate safety enforcement.

Information quality

  • Classification confidence: High
  • Reason for confidence: The reports provide highly detailed, quantitative results from a structured audit conducted by NewsGuard across 10 specific chatbots. The methodology, prompt counts, and specific false narratives are clearly documented, leaving little ambiguity about the AI systems' behavior.

An audit of ten leading AI chatbots revealed they repeated Russian disinformation narratives in 32% of responses, citing fake local news sites as reliable sources. This highlights how generative AI can be exploited to amplify coordinated foreign propaganda, posing a substantial threat to information security and democratic integrity during global election cycles.

  • Overall national security impact: Substantial
  • Response level: Substantial
  • Scope: Multiple nations
  • Primary target: United States
  • Other affected: Ukraine
  • Alleged perpetrator: Russia

Threat characteristics

  • Imminence: Near-term. The threat is highly relevant to ongoing global election cycles, requiring near-term monitoring and mitigation by security agencies and tech firms.
  • Autonomy: Human-supervised. The chatbots generate responses autonomously once prompted, but operate within guardrails and platforms managed by human developers.
  • Novelty: Evolved capability. Represents a significant evolution of traditional foreign disinformation campaigns, leveraging generative AI to automate and legitimize fake news.

Impact by dimension

  • Physical security: Negligible. No physical security threat or critical infrastructure compromise was reported in this incident.
  • Information security: Substantial. Ten leading AI chatbots repeated Russian disinformation in 32% of test prompts, citing fake local news sites. This demonstrates how AI can validate and amplify coordinated foreign propaganda campaigns at scale.
  • Sovereignty: Minor. The disinformation targeted democratic processes and U.S. election narratives, posing a minor threat to government trust, though no direct compromise of government systems occurred.
  • Economic security: Negligible. No direct threats to financial systems, strategic technologies, or economic infrastructure were identified.
  • Societal stability: Minor. The dissemination of false narratives could erode public trust and cause epistemic harm, but did not result in active civil unrest or mass human rights violations.
Explore in the interactive Incident Tracker