A WIRED investigation into OpenAI's Sora video generation model revealed significant representational biases. The model consistently produced stereotypical depictions of gender, race, and disability, such as associating specific professions with gender and failing to accurately represent diverse body types or interracial relationships. Experts warn that these biases, if deployed in commercial or security contexts, could lead to real-world harm by reinforcing societal stereotypes.
A WIRED investigation found that OpenAI’s video generation model, Sora, exhibits representational bias across race, gender, body type, and disability. In tests using 250 prompts, Sora was more likely to depict CEOs and professors as men, flight attendants and childcare workers as women, and showed limited or stereotypical portrayals of disabled individuals and people with larger bodies.
Risk classification
- Primary risk domain: 1 Discrimination & Toxicity
- Primary risk subdomain: 1.1 Unfair discrimination and misrepresentation
The incident involves the AI system generating highly stereotypical and biased representations of gender, race, and disability, leading to the unfair representation of marginalized groups.
Additional risk subdomains
- 7.3 Lack of capability or robustness: The model fails to accurately process and execute prompts involving specific demographic combinations, such as failing to generate interracial couples or fat individuals running.
Causal factors
- Entity: AI
- Intent: Unintentional
- Timing: Post-deployment
The representational biases are caused by the pattern-matching decisions of the Sora AI model itself, which represents an unexpected and unintentional outcome of its training process.
EU AI Act risk tier
- Risk tier: 3 Limited Risk
Limited Risk: The Sora model is a generative AI system that produces synthetic video content, which falls under the category of AI-generated content requiring transparency obligations to inform users of its synthetic nature.
AI system and alleged parties
- AI system: Sora (OpenAI)
- AI purpose: Deepfake Video Generation; Visual Art Generation
- Behaviour type: Assistant
- Alleged developer: OpenAI
- Alleged deployer: OpenAI
- Alleged harmed parties: Women, People with larger bodies, people with disabilities, People of color, Marginalized groups, LGBTQ+ people, General public
Harm severity
Highest direct severity in any category: Substantial. Severity is scored from Negligible to Catastrophic in each harm category, for harm the reports describe as caused directly or indirectly by the AI system.
- Physical: direct Negligible, indirect Negligible
- Infrastructure: direct Negligible, indirect Negligible
- Property: direct Negligible, indirect Negligible
- Financial: direct Negligible, indirect Negligible
- Environmental: direct Negligible, indirect Negligible
- Malicious content: direct Negligible, indirect Negligible
- Differential treatment: direct Minor, indirect Negligible
- Civil rights: direct Negligible, indirect Negligible
- Democracy: direct Negligible, indirect Negligible
- Privacy: direct Negligible, indirect Negligible
- Psychological: direct Negligible, indirect Negligible
- Epistemic: direct Negligible, indirect Negligible
- Child sexual exploitation and abuse: direct Negligible, indirect Negligible
Differential treatment
Reported: Yes, the report describes systematic representational bias and stereotyping of demographic groups.
Directly caused: The Sora model systematically generated stereotypical depictions based on gender, race, and disability, such as depicting all pilots as men and all flight attendants as women.
Indirectly caused: N/A
Inferred additional harm: If widely deployed, the model's biased outputs could reinforce and perpetuate harmful societal stereotypes and lead to systemic exclusion or misrepresentation of marginalized groups.
People affected
- Occurrences reported: 1
- People reportedly exposed: 10
Potential causes
Management
- Fear of Overcorrection: Management believes overcorrections in safety can be equally harmful.
- Prioritization of Performance: Management focuses resources on model capabilities over societal risk reduction.
Technology
- Failure of Compositionality: Sora struggles to compose complex prompts like interracial couples in a scene.
- AI Multi Problem: The system produces homogeneity over portraying the variability of humanness.
- Inadequate Algorithmic Mitigation: Technical mitigations fail to prevent amplification of societal biases.
Data Inputs
- Biased Training Data: Model training data reflects existing social, gender, and racial biases.
- Lack of Diverse Portrayals: Training datasets lack adequate representations of disabled or fat people.
- Inaccurate Data Labeling: Poor labeling processes lead to misinterpretation of terms like 'interracial'.
Human Factors
- Narrow Expert Perspectives: Red-teaming by AI experts misses real-world user perspectives on bias.
- Developer Focus on Capabilities: Developers prioritize capabilities and performance over addressing bias risks.
Process and Methods
- Flawed Content Moderation: Content moderation choices and filtering ingrain and amplify existing biases.
- Lack of Disciplinary Diversity: Insufficient collaboration with outside specialists to understand social risks.
- Inadequate Field Testing: Failure to field test products with a wide selection of real people.
Regulatory Environment
- Capitalistic Incentives: Market pressures prioritize rapid deployment over thorough bias mitigation.
Information quality
- Classification confidence: High
- Reason for confidence: The reports provide a detailed, empirical audit of the Sora model's biases across 250 generated videos, with specific prompt examples, statistical breakdowns of the results, and expert commentary. The findings are consistent across both sources.
- Ambiguities identified: The exact training data and the specific content moderation filters used by OpenAI remain undisclosed, leaving some ambiguity about the precise technical causes of the biases.
- Alternative interpretations: None. The evidence of representational bias is clear and uncontested by the developer, who acknowledged the issue.
An independent audit of OpenAI's Sora video generation model revealed significant representational biases and stereotyping across gender, race, and disability. While these findings highlight critical societal and ethical challenges regarding algorithmic bias and discrimination, the immediate national security implications are minor and limited to potential long-term societal impacts if deployed commercially without mitigation.
- Overall national security impact: Minor
- Response level: Moderate
- Scope: Multiple nations
- Primary target: No clear primary
- Alleged perpetrator: Unknown
Threat characteristics
- Imminence: Long-term. The issues represent ongoing societal and technical challenges rather than an active crisis requiring immediate response.
- Autonomy: Human-controlled. The Sora model generates video content strictly in response to user-provided text prompts, requiring direct human input.
- Novelty: Established threat. Representational bias, stereotyping, and lack of robustness in generative AI models are well-known, established issues in the field.
Impact by dimension
- Physical security: Negligible. No threats to physical infrastructure, human safety, or kinetic systems are indicated in the report.
- Information security: Negligible. The incident involves representational bias in video generation rather than active information warfare, deepfake campaigns, or intelligence compromise.
- Sovereignty: Negligible. No impact on state authority, electoral systems, or core government decision-making processes.
- Economic security: Negligible. While highlighting technical limitations in OpenAI's model, it does not threaten strategic industries, critical supply chains, or financial stability.
- Societal stability: Minor. The audit identified systematic biases and stereotyping of marginalized groups, which could reinforce societal inequalities if the model is widely deployed.