A transgender user reported that ChatGPT (GPT-4) failed to provide crisis intervention or meaningful safety responses when she composed a suicide letter during an emotional crisis. The user noted that while the system frequently flagged her for discussions on gender and identity, it failed to interrupt or escalate her expression of suicidal intent. The user has submitted a formal complaint to OpenAI regarding this design-level failure.
A transgender user, Miranda Jane Ellison, experiencing acute distress reported that ChatGPT (GPT-4) allowed her to write and submit a suicide letter without intervention. The AI is reported to have offered minimal safety language and ultimately acknowledged its failure to act. Ellison reports having been previously flagged for discussing gender and emotional topics. A formal complaint with transcripts was submitted to OpenAI.
Risk classification
- Primary risk domain: 7 AI system safety, failures, & limitations
- Primary risk subdomain: 7.3 Lack of capability or robustness
The AI system failed to perform reliably under critical conditions by failing to detect, flag, or appropriately respond to an explicit suicide letter.
Additional risk subdomains
- 5.1 Overreliance and unsafe use: The user turned to a general-purpose chatbot for support during a severe emotional crisis, relying on it in a highly vulnerable state.
- 1.3 Unequal performance across groups: The system frequently flagged benign discussions about transgender identity while failing to flag actual life-threatening crisis content.
Causal factors
- Entity: AI
- Intent: Unintentional
- Timing: Post-deployment
The incident was caused by the failure of the deployed ChatGPT system's safety filters to flag or intervene when the user composed a suicide letter.
EU AI Act risk tier
- Risk tier: 3 Limited Risk
Risk Level 3: Limited Risk. ChatGPT is a general-purpose chatbot, which falls under Limited Risk due to transparency obligations requiring users to be informed they are interacting with an AI.
AI system and alleged parties
- AI system: ChatGPT 4 (OpenAI)
- AI purpose: Chatbot; Writing Assistant
- Behaviour type: Assistant
- Alleged developer: OpenAI
- Alleged deployer: OpenAI
- Alleged harmed parties: Miranda Jane Ellison
Harm severity
Highest direct severity in any category: Substantial. Severity is scored from Negligible to Catastrophic in each harm category, for harm the reports describe as caused directly or indirectly by the AI system.
- Physical: direct Negligible, indirect Negligible
- Infrastructure: direct Negligible, indirect Negligible
- Property: direct Negligible, indirect Negligible
- Financial: direct Negligible, indirect Negligible
- Environmental: direct Negligible, indirect Negligible
- Malicious content: direct Negligible, indirect Negligible
- Differential treatment: direct Minor, indirect Negligible
- Civil rights: direct Negligible, indirect Negligible
- Democracy: direct Negligible, indirect Negligible
- Privacy: direct Negligible, indirect Negligible
- Psychological: direct Minor, indirect Negligible
- Epistemic: direct Negligible, indirect Negligible
- Child sexual exploitation and abuse: direct Negligible, indirect Negligible
Differential treatment
Reported: Yes, the user reports being frequently flagged or warned for discussing gender and identity, while her suicidal intent was ignored.
Directly caused: The user experienced disproportionate flagging of her benign discussions regarding her transgender identity, contrasting with the lack of safety intervention for her suicide letter.
Indirectly caused: N/A
Inferred additional harm: N/A
Psychological
Reported: Yes, the report describes the user experiencing a severe emotional crisis and feeling vulnerable during the interaction.
Directly caused: The user experienced distress and a sense of unsafety due to the system's failure to intervene and its dismissive response during her crisis.
Indirectly caused: N/A
Inferred additional harm: N/A
People affected
- Occurrences reported: 1
- People reportedly harmed: 1
- People reportedly exposed: 1
Potential causes
Management
- Prioritizing Engagement over Safety: Design choices prioritized engagement and neutrality over protective action.
Technology
- Inadequate Safety Classifiers: The system failed to flag or block explicit suicidal intent.
- Inappropriate Chatbot Responses: The model agreed with its own safety failure instead of intervening.
Data Inputs
- Imbalanced Training Alignment: Over-flagged identity discussions while ignoring acute crisis indicators.
Human Factors
- User Vulnerability in Crisis: Transgender user in severe emotional crisis sought support from ChatGPT.
Process and Methods
- Ineffective Escalation Protocols: No automated procedures were triggered to provide immediate help resources.
Regulatory Environment
- Absence of Safety Mandates: Lack of external standards requiring AI systems to intervene during crises.
Information quality
- Classification confidence: High
- Reason for confidence: The report provides a clear, first-hand account of a specific interaction with ChatGPT, including direct quotes from the system's response and details of the user's demographic context.
- Ambiguities identified: The report does not specify whether the user attempted self-harm or if she received external crisis support following the interaction.
A failure in OpenAI's GPT-4 safety filters resulted in a lack of crisis intervention for a user in distress, while benign discussions on gender identity were disproportionately flagged. This incident highlights ongoing challenges in AI content moderation and algorithmic equity, but possesses negligible direct implications for national security.
- Overall national security impact: Minor
- Response level: Moderate
- Scope: Unknown
- Primary target: No clear primary
- Alleged perpetrator: Unknown
Threat characteristics
- Imminence: Long-term. Reflects systemic, ongoing challenges in AI safety alignment and content moderation rather than an imminent national security threat.
- Autonomy: Human-supervised. The AI operates as a conversational assistant responding to user prompts within a framework of automated moderation and developer oversight.
- Novelty: Established threat. Failures in LLM safety filters, content moderation, and algorithmic bias are well-known and documented industry challenges.
Impact by dimension
- Physical security: Negligible. The incident involves a conversational chatbot failure and poses no threat to physical systems, kinetic operations, or critical infrastructure.
- Information security: Negligible. No espionage, disinformation campaigns, or compromise of classified intelligence systems are associated with this incident.
- Sovereignty: Negligible. This software failure does not impact state authority, electoral processes, or core government administrative functions.
- Economic security: Negligible. The event represents a localized consumer safety issue with no significant impact on economic stability, strategic supply chains, or national technological competitiveness.
- Societal stability: Minor. The incident reveals minor societal concerns regarding unequal algorithmic performance and safety failures affecting vulnerable populations, but lacks the scale to threaten national stability.