Researchers at Cornell University discovered that OpenAI's Whisper speech-to-text system frequently hallucinates violent language, fake websites, and fabricated personal information when processing audio containing long pauses. These hallucinations disproportionately affect individuals with speech impairments, such as those with aphasia. The study highlights significant risks for downstream applications in sensitive fields like legal, medical, and hiring processes.
Researchers at Cornell reportedly found that OpenAI's Whisper, a speech-to-text system, can hallucinate violent language and fabricated details, especially with long pauses in speech, such as from those with speech impairments. Analyzing 13,000 clips, they determined 1% contained harmful hallucinations. These errors pose risks in hiring, legal trials, and medical documentation. The study suggests improving model training to reduce these hallucinations for diverse speaking patterns.
Risk classification
- Primary risk domain: 1 Discrimination & Toxicity
- Primary risk subdomain: 1.3 Unequal performance across groups
Whisper performs unequally across groups because it is significantly more likely to hallucinate and produce errors when transcribing speech from individuals with speech impairments due to their longer pauses.
Additional risk subdomains
- 7.3 Lack of capability or robustness: The system fails to perform robustly under varying conditions, specifically when encountering silences or pauses in audio, leading to hallucinations.
- 1.2 Exposure to toxic content: The system hallucinates violent language, potentially exposing users or downstream readers to harmful and inappropriate content.
Causal factors
- Entity: AI
- Intent: Unintentional
- Timing: Post-deployment
The risk is caused by Whisper's algorithmic hallucinations, which are an unexpected and unintentional outcome of the system's deployment.
EU AI Act risk tier
High Risk: The report describes Whisper's potential downstream use in high-risk areas such as 'AI-based hiring, in courtroom trials or patient notes in medical settings', which fall under the High Risk (Level 2) classification of the EU AI Act due to their significant implications for employment, justice, and healthcare.
AI system and alleged parties
- AI system: Whisper (OpenAI)
- AI purpose: Voice Recognition
- Behaviour type: Tool
- Alleged developer: OpenAI
- Alleged deployer: Whisper, Organizations integrating Whisper into customer service systems, OpenAI, Companies using Whisper
- Alleged harmed parties: Users whose speech is misinterpreted by Whisper, Professionals relying on accurate transcriptions, Individuals with speech impairments, General public
Harm severity
Highest direct severity in any category: Substantial. Severity is scored from Negligible to Catastrophic in each harm category, for harm the reports describe as caused directly or indirectly by the AI system.
- Physical: direct Negligible, indirect Negligible
- Infrastructure: direct Negligible, indirect Negligible
- Property: direct Negligible, indirect Negligible
- Financial: direct Negligible, indirect Negligible
- Environmental: direct Negligible, indirect Negligible
- Malicious content: direct Negligible, indirect Negligible
- Differential treatment: direct Minor, indirect Minor
- Civil rights: direct Negligible, indirect Negligible
- Democracy: direct Negligible, indirect Negligible
- Privacy: direct Negligible, indirect Negligible
- Psychological: direct Negligible, indirect Negligible
- Epistemic: direct Negligible, indirect Negligible
- Child sexual exploitation and abuse: direct Negligible, indirect Negligible
Differential treatment
Reported: Yes, the report notes that Whisper is more likely to hallucinate when analyzing speech from people with speech impairments, such as aphasia.
Directly caused: People with speech impairments experience significantly higher rates of transcription errors and harmful hallucinations compared to those without impairments.
Indirectly caused: This could lead to unfair outcomes if the transcripts are used in downstream applications like hiring, court trials, or medical settings.
Inferred additional harm: Potential systemic exclusion or misrepresentation of speech-impaired individuals in professional or legal settings relying on automated transcriptions.
People affected
- Occurrences reported: 1
- People reportedly exposed: 1000
Potential causes
Technology
- Silence treated as words: The underlying LLM technology interprets silence or pauses as words.
- Sensitivity to speech pauses: Whisper is more likely to hallucinate when speech has long pauses.
Data Inputs
- Lack of diverse speech training: Training data did not adequately represent speakers with speech impairments.
Human Factors
- User speech impairments: Aphasia and other conditions cause longer pauses, triggering hallucinations.
Process and Methods
- Flawed modeling choices: Modeling choices failed to account for nonvocal durations during speech.
Information quality
- Classification confidence: High
- Reason for confidence: The report is based on peer-reviewed research presented at a major conference (FAccT) by Cornell University researchers. It provides specific statistics, such as a 1% hallucination rate across 13,000 tested clips, and clearly identifies the mechanism (long pauses triggering hallucinations). There is high certainty about the technical behavior and the affected demographic group.
- Ambiguities identified: The exact number of unique individuals represented in the 13,000 AphasiaBank clips is not specified.
- Alternative interpretations: None. The technical failure and its disproportionate impact on speech-impaired individuals are clearly documented.
Cornell researchers found that OpenAI Whisper speech-to-text system frequently hallucinates violent language and fake information when processing audio with long pauses, disproportionately affecting speech-impaired individuals. While presenting safety and discrimination risks for downstream legal or medical use, the direct national security impact remains negligible.
- Overall national security impact: Minor
- Response level: Moderate
- Scope: Multiple nations
- Primary target: No clear primary
- Other affected: Unknown
- Alleged perpetrator: Unknown
Threat characteristics
- Imminence: Long-term. This is an ongoing technical robustness issue rather than an active crisis requiring immediate national security response.
- Autonomy: Human-controlled. Whisper operates as a transcription tool assisting humans, who ultimately control the input and review the output.
- Novelty: Established threat. AI model hallucinations and robustness issues are well-established phenomena, though the specific trigger of speech pauses is a newly detailed vulnerability.
Impact by dimension
- Physical security: Negligible. No physical systems, infrastructure, or human safety threats were reported or impacted in this incident.
- Information security: Negligible. The incident involves unintentional model hallucinations during transcription, not coordinated information warfare or compromise of classified intelligence.
- Sovereignty: Negligible. No actual compromise of government functions, sovereignty, or decision-making processes was reported in the incident details.
- Economic security: Negligible. The incident represents a software robustness issue and does not involve strategic technology theft or threats to economic stability.
- Societal stability: Minor. The transcription errors disproportionately affect individuals with speech impairments, presenting potential discrimination and unequal performance risks in downstream applications.