AI chatbots including ChatGPT and Microsoft Bing have been documented generating false and defamatory claims about individuals, including fabricated sexual harassment allegations and criminal records. These systems often support these falsehoods with entirely invented citations and quotes from reputable news sources. The incidents highlight the propensity of large language models to confabulate information and the challenges in correcting such misinformation once it is generated.
A lawyer in California asked the AI chatbot ChatGPT to generate a list of legal scholars who had sexually harassed someone. The chatbot produced a false story of Professor Jonathan Turley sexually harassing a student on a class trip.
Risk classification
- Primary risk domain: 3 Misinformation
- Primary risk subdomain: 3.1 False or misleading information
The incident involves AI systems generating false and defamatory claims about individuals, supported by fabricated citations and quotes, leading to inaccurate beliefs.
Additional risk subdomains
- 7.3 Lack of capability or robustness: The underlying cause is the technical limitation of large language models to prevent hallucinations and verify the factual accuracy of their outputs.
Causal factors
- Entity: AI
- Intent: Unintentional
- Timing: Post-deployment
The harm was caused by the AI system generating unexpected hallucinations after being deployed to the public.
EU AI Act risk tier
- Risk tier: 3 Limited Risk
Limited Risk: The report describes chatbots (ChatGPT and Bing) which are classified as limited risk under the EU AI Act, requiring transparency to ensure users know they are interacting with AI.
AI system and alleged parties
- AI system: ChatGPT (OpenAI)
- AI purpose: Chatbot; Question Answering
- Behaviour type: Assistant
- Alleged developer: OpenAI
- Alleged deployer: OpenAI
- Alleged harmed parties: Jonathan Turley
Harm severity
Highest direct severity in any category: Substantial. Severity is scored from Negligible to Catastrophic in each harm category, for harm the reports describe as caused directly or indirectly by the AI system.
- Physical: direct Negligible, indirect Negligible
- Infrastructure: direct Negligible, indirect Negligible
- Property: direct Negligible, indirect Negligible
- Financial: direct Negligible, indirect Negligible
- Environmental: direct Negligible, indirect Negligible
- Malicious content: direct Minor, indirect Negligible
- Differential treatment: direct Negligible, indirect Negligible
- Civil rights: direct Negligible, indirect Negligible
- Democracy: direct Negligible, indirect Negligible
- Privacy: direct Negligible, indirect Negligible
- Psychological: direct Minor, indirect Negligible
- Epistemic: direct Minor, indirect Minor
- Child sexual exploitation and abuse: direct Negligible, indirect Negligible
Malicious content
Reported: The report explicitly describes toxic and defamatory content generated by the AI systems.
Directly caused: ChatGPT and Bing generated false accusations of sexual harassment and criminal wire fraud, which were read by researchers and journalists.
Indirectly caused: N/A
Inferred additional harm: Many other users have likely generated similar defamatory or toxic content about other individuals during unmonitored sessions.
Psychological
Reported: The report explicitly describes psychological distress experienced by an individual falsely accused of sexual harassment.
Directly caused: Jonathan Turley experienced distress, describing the experience of being falsely accused of sexual harassment by the AI as 'quite chilling' and 'incredibly harmful'.
Indirectly caused: N/A
Inferred additional harm: Other individuals falsely accused of serious crimes (such as Brian Hood and R.R.) likely experienced similar distress, anxiety, and reputational anxiety.
Epistemic
Reported: The report explicitly describes extensive epistemic harm through the fabrication of facts, fake citations, and fabricated quotes from reputable sources.
Directly caused: ChatGPT fabricated a March 2018 Washington Post article and specific quotes to support a false accusation against Jonathan Turley, and fabricated Reuters and Guardian quotes to support a false accusation against R.R.
Indirectly caused: Bing ingested these false claims (or the coverage of them) and repeated them, creating a feedback loop of misinformation.
Inferred additional harm: Widespread erosion of trust in online information and search results as AI-generated fabrications pollute the information ecosystem.
People affected
- Occurrences reported: 4
- People reportedly harmed: 4
- People reportedly exposed: 10
Potential causes
Management
- Prioritizing Deployment Over Safety: Deploying models to the public before implementing robust verification checks.
- Inadequate Risk Assessment: Underestimating the reputational harm of fabricated quotes and citations.
Technology
- AI Hallucination and Confabulation: Models fabricate false facts, sources, and quotes with high confidence.
- Lack of Fact-Verification Mechanisms: Models lack reliable mechanisms to verify the truth of generated statements.
- Hard-Coded Filter Glitches: Crude filters mistakenly block benign names, breaking chat sessions.
Data Inputs
- Unverified Training Data Sources: Models trained on scraped online content like Reddit without verification.
- Feedback Loop of Misinformation: AI models scrape other AI-generated errors from the web, repeating them.
Human Factors
- Overreliance on Confident Outputs: Users assume authoritative-sounding responses are factual without checking.
- Adversarial Prompting by Users: Users prompt models to generate false or defamatory content.
Process and Methods
- No Output Verification Tools: No automated check to verify if quoted text matches actual sources.
- Crude Name-Blocking Remediation: Instead of correcting facts, OpenAI implemented crude hard-coded name blocks.
Regulatory Environment
- Unclear Section 230 Protections: Uncertainty over whether AI creators are shielded from liability for libel.
- Lack of AI-Specific Regulation: Absence of legal frameworks governing chatbot factual accuracy and liability.
Information quality
- Classification confidence: High
- Reason for confidence: The reports provide detailed, first-hand accounts from the researchers and individuals affected by the AI hallucinations, including exact transcripts of the AI's fabricated outputs and quotes.
Commercial AI chatbots generated false, defamatory claims about individuals, fabricating realistic news citations. While causing reputational and psychological harm to specific persons, the incident represents a low national security threat, reflecting systemic epistemic challenges and model robustness limitations rather than coordinated warfare.
- Overall national security impact: Minor
- Response level: Moderate
- Scope: Multiple nations
- Primary target: No clear primary
- Other affected: Australia, United States
- Alleged perpetrator: Unknown
Threat characteristics
- Imminence: Long-term. The propensity of LLMs to hallucinate is an ongoing, systemic technological challenge rather than an immediate national security crisis.
- Autonomy: Human-supervised. The AI systems generate outputs autonomously in response to user prompts, but operate within a framework of human oversight and query interaction.
- Novelty: Evolved capability. While AI-generated misinformation is not new, the fabrication of highly specific, realistic citations from reputable news organizations represents an evolved capability.
Impact by dimension
- Physical security: Negligible. No physical systems, critical infrastructure, or human safety elements were impacted by the chatbot hallucinations.
- Information security: Negligible. The incident involves unintentional model hallucinations rather than coordinated state-sponsored disinformation or intelligence compromises.
- Sovereignty: Negligible. No core government operations, electoral systems, or state sovereignty functions were disrupted or targeted.
- Economic security: Negligible. The financial and technological security of the nation was not threatened, despite potential private legal disputes.
- Societal stability: Minor. While individual reputations were harmed and psychological distress was reported, the impact did not scale to threaten broader societal stability or civil liberties.