Following an early leak of the open-source model Stable Diffusion on 4chan, users utilized the system to generate non-consensual pornographic deepfakes of celebrities. Experts and researchers highlighted that the model's lack of immutable safety controls, combined with its accessibility, enables the creation of harmful, targeted content at scale, posing risks of harassment, extortion, and the exacerbation of existing illegal behaviors.
Stable Diffusion, an open-source image generation model by Stability AI, was reportedly leaked on 4chan prior to its release date, and was used by its users to generate pornographic deepfakes of celebrities.
Risk classification
- Primary risk domain: 4 Malicious actors
- Primary risk subdomain: 4.3 Fraud, scams, and targeted manipulation
The incident primarily involves the creation of non-consensual pornographic deepfakes and targeted sexual imagery to humiliate, harass, or potentially blackmail specific individuals.
Additional risk subdomains
- 1.2 Exposure to toxic content: The leaked model was used on public forums like 4chan to generate and distribute explicit pornographic and sexually explicit content.
Causal factors
- Entity: Human
- Intent: Intentional
- Timing: Post-deployment
The primary risk and harm are caused by human users intentionally prompting the leaked, open-source model to generate non-consensual pornographic deepfakes.
EU AI Act risk tier
- Risk tier: 3 Limited Risk
Limited Risk: The report describes an AI system used to generate synthetic content and deepfakes, which falls under transparency and disclosure obligations under the EU AI Act.
AI system and alleged parties
- AI system: Safety Classifier, Stable Diffusion (Stability AI)
- AI purpose: Deepfake Image Generation; Visual Art Generation
- Behaviour type: Tool
- Alleged developer: Stability AI, Runway, LAION, EleutherAI, CompVis LMU
- Alleged deployer: Stability AI
- Alleged harmed parties: Stability AI, deepfaked celebrities
Harm severity
Highest direct severity in any category: Substantial. Severity is scored from Negligible to Catastrophic in each harm category, for harm the reports describe as caused directly or indirectly by the AI system.
- Physical: direct Negligible, indirect Negligible
- Infrastructure: direct Negligible, indirect Negligible
- Property: direct Negligible, indirect Negligible
- Financial: direct Negligible, indirect Negligible
- Environmental: direct Negligible, indirect Negligible
- Malicious content: direct Minor, indirect Substantial
- Differential treatment: direct Negligible, indirect Minor
- Civil rights: direct Negligible, indirect Minor
- Democracy: direct Negligible, indirect Negligible
- Privacy: direct Negligible, indirect Minor
- Psychological: direct Negligible, indirect Minor
- Epistemic: direct Minor, indirect Minor
- Child sexual exploitation and abuse: direct Negligible, indirect Negligible
Malicious content
Reported: The report explicitly describes the generation and spread of toxic and malicious content, specifically non-consensual pornographic deepfakes.
Directly caused: Users on 4chan utilized the leaked Stable Diffusion model to directly generate explicit, nude images of celebrities.
Indirectly caused: The generated pornographic images were shared across online discussion boards, exposing forum users to non-consensual explicit content.
Inferred additional harm: It is highly likely that thousands of additional explicit deepfake images have been generated and distributed across various adult websites and forums.
Differential treatment
Reported: The report explicitly describes differential treatment, noting that women are disproportionately targeted by non-consensual deepfakes.
Directly caused: N/A
Indirectly caused: The proliferation of easy-to-use deepfake tools disproportionately subjects women and female celebrities to sexual exploitation and harassment online.
Inferred additional harm: The widespread availability of these tools likely exacerbates gender-based harassment and unequal safety outcomes for women on social media platforms.
Civil rights
Reported: The report explicitly describes violations of human and civil rights, characterizing the deepfakes as sexual exploitation and harassment.
Directly caused: N/A
Indirectly caused: The generation of non-consensual sexual imagery violates the victims' rights to personal dignity, privacy, and freedom from gender-based harassment.
Inferred additional harm: Widespread, unmonitored generation of deepfakes likely leads to systemic violations of bodily autonomy and privacy rights for many targeted individuals.
Privacy
Reported: The report explicitly describes privacy concerns, discussing how personal photos scraped from social media can be used to condition models to generate targeted pornographic imagery.
Directly caused: N/A
Indirectly caused: The public likenesses and personal photos of individuals were exploited without consent to generate intimate, explicit content.
Inferred additional harm: It is highly likely that many private individuals have had their personal social media photos weaponized to generate non-consensual explicit imagery, violating their privacy.
Psychological
Reported: The report explicitly describes psychological harm, noting that the creation of objectionable content in a victim's likeness causes real trauma and severe harassment.
Directly caused: N/A
Indirectly caused: Victims of deepfake campaigns, such as journalist Rana Ayyub, experienced severe online harassment and trauma, requiring international intervention.
Inferred additional harm: It is highly likely that numerous unnamed celebrities and private individuals whose likenesses were deepfaked on 4chan suffered severe distress, anxiety, and psychological trauma.
Epistemic
Reported: The report explicitly describes epistemic harm, noting the risk of creating fake but potentially damaging footage that could implicate someone in a crime they did not commit.
Directly caused: The AI system generated highly realistic, fabricated images of real individuals.
Indirectly caused: The dissemination of photorealistic fakes misleads viewers and undermines the perceived reliability of photographic evidence.
Inferred additional harm: The ease of generating convincing fakes likely contributes to a broader erosion of shared truth and trust in digital media.
People affected
- Occurrences reported: 1
- People reportedly harmed: 5
- People reportedly exposed: 5
Potential causes
Management
- Open Source Release Strategy: Releasing the model openly removed the ability to monitor or block abuse.
- Inadequate Risk Assessment: Management did not fully foresee the scale of automated blackmail attacks.
- Lack of Pre-release Governance: Inadequate internal coordination before the model leaked on public forums.
Technology
- Bypassable Safety Classifier: The built-in Safety Classifier is optional and can be easily disabled.
- Unfettered Model Architecture: The model itself is not restricted at a technical level from generating porn.
- Low Consumer Hardware Barriers: Model runs on cheap consumer GPUs, enabling offline, unmonitored generation.
Data Inputs
- Unfiltered Training Data: Model was trained on datasets allowing generation of celebrity likenesses.
- Scraped Social Media Photos: Personal photos can be used to condition the model for targeted deepfakes.
- Lack of Consent for Training: Training data included celebrity likenesses without explicit consent.
Human Factors
- Malicious User Intent: Bad actors on forums like 4chan actively seek to generate non-consensual porn.
- Underestimating Harmful Uses: Developers underappreciated how the system could be used in unanticipated ways.
- Harassment and Extortion: Trolls and blackmailers utilize deepfakes to target and exploit individuals.
Process and Methods
- Lack of Post-Deployment Controls: Once released in the wild, API rate limits and safety controls are bypassed.
- Early Model Leakage: The model leaked early on 4chan before official safety gates were finalized.
- Ineffective License Enforcement: The open license prohibited abuse but lacked technical enforcement mechanisms.
Regulatory Environment
- Absence of Deepfake Laws: Few laws specifically protect individuals against AI-generated pornography.
- Uneven Platform Enforcement: Image hosts and platforms enforce policies against deepfakes inconsistently.
- Lack of Global Standards: No global standards exist to regulate open-source generative AI deployment.
Information quality
- Classification confidence: High
- Reason for confidence: The report provides clear, detailed accounts of the Stable Diffusion leak, the specific platform where it occurred (4chan), the nature of the generated content (celebrity deepfakes), and the societal risks identified by experts. The roles of the developer and the users are explicitly defined.
- Ambiguities identified: None of significance; the distinction between the general capabilities of the model and the specific malicious actions of the users is clear.
- Alternative interpretations: The incident could be viewed purely as a copyright or open-source licensing issue, but the primary focus of the reported harm is clearly on malicious use and targeted harassment.
The leak and subsequent abuse of the open-source Stable Diffusion model to generate non-consensual celebrity deepfakes highlights the evolving threat of accessible synthetic media. While posing minimal direct threat to critical infrastructure or state sovereignty, the incident demonstrates how unchecked generative AI capabilities can scale societal harms, privacy violations, and targeted harassment.
- Overall national security impact: Minor
- Response level: Moderate
- Scope: Multiple nations
- Primary target: No clear primary
- Other affected: Global
- Alleged perpetrator: 4chan users
Threat characteristics
- Imminence: Long-term. The proliferation of open-source generative models represents an ongoing, long-term strategic challenge rather than an immediate national security crisis.
- Autonomy: Human-controlled. The AI functions purely as a tool, generating images directly from user-provided text prompts without independent decision-making.
- Novelty: Evolved capability. Deepfakes are an established threat, but the release of a high-fidelity, open-source model with easily bypassed safety filters represents a significant advancement in accessibility.
Impact by dimension
- Physical security: Negligible. The incident involves image generation software and does not pose any kinetic threats to physical systems, human safety, or critical infrastructure.
- Information security: Minor. The technology enables high-fidelity synthetic media that can be used for targeted blackmail, extortion, and discrediting public figures, though this incident was primarily focused on harassment.
- Sovereignty: Negligible. No state authority, electoral systems, or core government decision-making processes were compromised or targeted in this incident.
- Economic security: Negligible. While the open-source model was leaked, it does not represent a critical threat to national economic stability or strategic technological defense advantages.
- Societal stability: Minor. The widespread generation of non-consensual pornographic deepfakes disproportionately targets women, violating privacy rights and causing psychological harm, but remains manageable under standard legal frameworks.