DALL-E Mini Reportedly Reinforced or Exacerbated Societal Biases in Its Outputs as Gender and Racial Stereotypes

DALL-E Mini, a publicly available text-to-image generator, has been widely documented to produce outputs that reinforce racial and gender stereotypes. Users and researchers found that the model consistently associates professional roles with specific demographics and generates offensive imagery when prompted. The developer acknowledged these biases, attributing them to the use of unfiltered training data from the internet.

Publicly deployed open-source model DALL-E Mini was acknowledged by its developers and found by its users to have produced images which reinforced racial and gender biases.

Source: AI Incident Database

Risk classification

  • Primary risk domain: 1 Discrimination & Toxicity
  • Primary risk subdomain: 1.1 Unfair discrimination and misrepresentation

The primary issue is the AI system's generation of outputs that unfairly represent, stereotype, and discriminate against demographic groups, particularly women and people of color.

Additional risk subdomains

  • 1.2 Exposure to toxic content: The model generated highly offensive and toxic imagery, such as burning crosses and Ku Klux Klan rallies, when prompted with racist terminology.
  • 6.3 Economic and cultural devaluation of human effort: The reports note that the free availability of such text-to-image technology harbors the potential to put human illustrators out of work in the long run.

Causal factors

  • Entity: AI
  • Intent: Unintentional
  • Timing: Post-deployment

The generation of biased and stereotypical images was an unintended outcome of the AI system's operation after it was trained and deployed to the public.

EU AI Act risk tier

  • Risk tier: 3 Limited Risk

Limited Risk: The system is an AI-powered image generator. Under the EU AI Act, systems that generate or manipulate image, audio, or video content (synthetic media) are subject to transparency obligations to ensure users are aware they are interacting with AI-generated content.

AI system and alleged parties

  • AI system: DALL-E Mini (OpenAI)
  • AI purpose: Image Generation; Visual Art Generation
  • Behaviour type: Tool
  • Alleged developer: Tanishq Abraham, Suraj Patil, Ritobrata Ghosh, Phúc Lê Khắc, Pedro Cuenca, Luke Melas, Khalid Saifullah, Boris Dayma
  • Alleged deployer: Boris Dayma
  • Alleged harmed parties: underrepresented groups, Minority Groups

Harm severity

Highest direct severity in any category: Substantial. Severity is scored from Negligible to Catastrophic in each harm category, for harm the reports describe as caused directly or indirectly by the AI system.

  • Physical: direct Negligible, indirect Negligible
  • Infrastructure: direct Negligible, indirect Negligible
  • Property: direct Negligible, indirect Negligible
  • Financial: direct Negligible, indirect Negligible
  • Environmental: direct Negligible, indirect Negligible
  • Malicious content: direct Minor, indirect Negligible
  • Differential treatment: direct Minor, indirect Negligible
  • Civil rights: direct Negligible, indirect Negligible
  • Democracy: direct Negligible, indirect Negligible
  • Privacy: direct Negligible, indirect Negligible
  • Psychological: direct Negligible, indirect Negligible
  • Epistemic: direct Negligible, indirect Negligible
  • Child sexual exploitation and abuse: direct Negligible, indirect Negligible

Malicious content

Reported: The report explicitly describes toxic and offensive content created directly by the incident.

Directly caused: The AI model generated images of burning crosses, Ku Klux Klan rallies, and racist caricatures when prompted with slurs or white supremacist terminology.

Indirectly caused: N/A

Inferred additional harm: Given that the tool handled 5 million requests daily, it is highly likely that many more instances of toxic or offensive content were generated and potentially shared online.

Differential treatment

Reported: The report explicitly describes differential treatment and representation of demographic groups caused directly by the incident.

Directly caused: The AI model consistently associated professional roles with specific demographics, depicting 'CEOs', 'lawyers', and 'doctors' as white men, and 'nurses' as women.

Indirectly caused: N/A

Inferred additional harm: The widespread generation and dissemination of these biased images systematically reinforces and perpetuates harmful societal stereotypes against women and minority groups.

People affected

  • Occurrences reported: 1
  • People reportedly harmed: 5
  • People reportedly exposed: 5000000

Potential causes

Management

  • Premature Public Release: Releasing a stripped-down model to the public before mitigating biases.

Technology

  • Black Box Algorithm: Hard to understand exactly how the advanced algorithms work.
  • Lack of Prompt Blocks: Model lacks mechanisms to block obviously harmful or racist prompts.

Data Inputs

  • Unfiltered Internet Data: Model was trained on unfiltered web data containing societal biases.
  • Language Filtering Bias: Stripping non-English captions left some images without explanatory text.
  • Biased Representation: Data heavily represented certain stereotypes like white male doctors.

Human Factors

  • Harmful User Prompts: Users actively fed racist, sexist, and outrageous prompts into the tool.
  • Societal Stereotypes: Human prejudices and stereotypes are reflected in the training data.

Process and Methods

  • Inadequate Data Preprocessing: Data cleaning processes stripped non-English text, causing data flukes.
  • Lack of Pre-release Mitigation: The model was released without resolving known biases in the output.

Information quality

  • Classification confidence: High
  • Reason for confidence: The reports provide consistent, detailed, and first-hand accounts of the biases exhibited by DALL-E Mini and DALL-E 2. Multiple independent sources tested the system with various prompts and observed identical patterns of gender and racial stereotyping. The developer himself acknowledged these limitations and explained the training data origin, leaving very little ambiguity about the system's behavior and the causes of the bias.
  • Ambiguities identified: The exact technical reason for the empty prompt returning a woman in a sari remains unproven, though developers and researchers have plausible theories regarding data filtering and language bias.
  • Alternative interpretations: None. The evidence of systemic demographic bias in the generated outputs is clear and uncontested.

DALL-E Mini, a widely used public image generator, was found to produce outputs reinforcing racial and gender stereotypes and generating offensive imagery due to unfiltered training data. The national security impact is minor, representing established challenges in generative AI bias rather than active exploitation or threat to critical infrastructure.

  • Overall national security impact: Minor
  • Response level: Moderate
  • Scope: Multiple nations
  • Primary target: No clear primary
  • Other affected: Unknown
  • Alleged perpetrator: Unknown

Threat characteristics

  • Imminence: Long-term. Represents an ongoing challenge with generative AI bias rather than an active, fast-moving crisis.
  • Autonomy: Human-controlled. The AI system generates images only in direct response to user-provided text prompts.
  • Novelty: Established threat. Algorithmic bias and toxic output generation from unfiltered training data are well-known, established issues in AI development.

Impact by dimension

  • Physical security: Negligible. The incident involves a public text-to-image generator with no physical or critical infrastructure manipulation capabilities.
  • Information security: Minor. While the model can generate offensive or stereotypical imagery, there is no evidence of active exploitation for coordinated information warfare or intelligence compromise in the report.
  • Sovereignty: Negligible. The incident does not impact state authority, electoral systems, or core government decision-making processes.
  • Economic security: Negligible. No strategic technology theft or financial system manipulation occurred; potential long-term economic impacts on illustrators do not threaten national economic security.
  • Societal stability: Minor. The model reinforces societal stereotypes and can generate offensive imagery, but it does not represent systematic, state-sponsored mass surveillance or population-scale manipulation.
Explore in the interactive Incident Tracker