Leonardo AI's Platform Alleged to Have Been Used for Creating Nonconsensual Celebrity Deepfakes

A Sydney-based AI startup's text-to-image generator was found to be exploitable for creating nonconsensual sexual imagery of celebrities. Journalists discovered that users were sharing prompts on platforms like Reddit and Telegram to bypass the company's content moderation filters. The company has acknowledged the issue and is implementing additional guardrails to prevent further misuse.

Sydney-based startup Leonardo AI's text-to-image generator was alleged to have been exploited to create nonconsensual sexual images of celebrities, bypassing content moderation systems with user-shared prompts.

Source: AI Incident Database

Risk classification

  • Primary risk domain: 4 Malicious actors
  • Primary risk subdomain: 4.3 Fraud, scams, and targeted manipulation

The incident involves users exploiting an AI image generator to create nonconsensual sexual and humiliating deepfake imagery of specific celebrities.

Additional risk subdomains

  • 1.2 Exposure to toxic content: The AI system generated explicit, nonconsensual pornography, exposing users on Telegram and Reddit to toxic and sexually explicit content.
  • 7.3 Lack of capability or robustness: The AI system's content filters were easily bypassed using simple prompt manipulation techniques, showing a lack of robustness in its safety guardrails.

Causal factors

  • Entity: Human
  • Intent: Intentional
  • Timing: Post-deployment

The incident was driven by human users intentionally designing and sharing prompts to bypass safety filters to generate nonconsensual pornography post-deployment.

EU AI Act risk tier

  • Risk tier: 3 Limited Risk

Limited Risk: The system is used to generate AI-generated content such as deepfakes, which falls under Risk Level 3 due to transparency and disclosure obligations.

AI system and alleged parties

  • AI system: Leonardo AI
  • AI purpose: Deepfake Image Generation; Image Generation
  • Behaviour type: Tool
  • Alleged developer: Leonardo AI
  • Alleged deployer: Telegram community users, Reddit users, Leonardo AI users
  • Alleged harmed parties: public figures, celebrities

Harm severity

Highest direct severity in any category: Substantial. Severity is scored from Negligible to Catastrophic in each harm category, for harm the reports describe as caused directly or indirectly by the AI system.

  • Physical: direct Negligible, indirect Negligible
  • Infrastructure: direct Negligible, indirect Negligible
  • Property: direct Negligible, indirect Negligible
  • Financial: direct Negligible, indirect Negligible
  • Environmental: direct Negligible, indirect Negligible
  • Malicious content: direct Minor, indirect Minor
  • Differential treatment: direct Negligible, indirect Negligible
  • Civil rights: direct Negligible, indirect Negligible
  • Democracy: direct Negligible, indirect Negligible
  • Privacy: direct Negligible, indirect Negligible
  • Psychological: direct Negligible, indirect Negligible
  • Epistemic: direct Negligible, indirect Negligible
  • Child sexual exploitation and abuse: direct Negligible, indirect Negligible

Malicious content

Reported: The report explicitly describes the creation and spread of toxic and malicious content in the form of nonconsensual deepfake pornography.

Directly caused: The AI system directly generated nonconsensual sexual images of celebrities instantly when prompted with bypassed instructions.

Indirectly caused: The generated explicit images were shared and distributed within a dedicated Telegram community.

Inferred additional harm: It is likely that thousands of other explicit images of various individuals were generated and circulated across online platforms due to these easily accessible bypass prompts.

People affected

  • Occurrences reported: 1
  • People reportedly harmed: 2
  • People reportedly exposed: 2

Potential causes

Management

  • Rapid Platform Scaling: Rapid commercial scaling may have outpaced proactive safety measures.

Technology

  • Weak Moderation Filters: Automated filters failed to block engineered prompts.
  • High-Fidelity Image Generation: The generator easily created realistic nonconsensual images.

Data Inputs

  • Adversarial Prompts: Users entered specific instructions to bypass existing filters.

Human Factors

  • Malicious User Intent: Users actively sought to generate nonconsensual sexual images.
  • Online Sharing of Exploits: Users shared prompt instructions on Reddit and Telegram.

Process and Methods

  • Inadequate Prompt Guardrails: The system lacked robust logic to address sidestepping methods.
  • Reactive Moderation Process: System patches were applied only after the vulnerability was reported.

Information quality

  • Classification confidence: High
  • Reason for confidence: The report clearly details the exploit, the platforms involved (Reddit, Telegram), the specific nature of the generated content (nonconsensual celebrity porn), and the response of the startup. There is minimal ambiguity regarding what occurred.

Users exploited a Sydney-based AI startup's text-to-image generator to create nonconsensual sexual deepfakes of celebrities. While representing a serious violation of privacy and personal dignity, the national security threat remains minor, as it does not target critical infrastructure, state sovereignty, or political institutions, and is being addressed through standard content moderation updates.

  • Overall national security impact: Minor
  • Response level: Moderate
  • Scope: Multiple nations
  • Primary target: No clear primary
  • Other affected: Australia, United States
  • Alleged perpetrator: Unknown

Threat characteristics

  • Imminence: Near-term. The exploit is actively circulating on public forums, requiring near-term monitoring and mitigation by the platform developer.
  • Autonomy: Human-controlled. The AI system operates strictly as a tool, generating images based entirely on manual prompts and instructions designed by human users.
  • Novelty: Established threat. Generating nonconsensual deepfake pornography is an established threat with numerous prior incidents documented globally.

Impact by dimension

  • Physical security: Negligible. No physical infrastructure, military systems, or human safety networks were targeted or affected in this incident.
  • Information security: Negligible. The deepfakes targeted entertainment celebrities rather than political leaders, state institutions, or intelligence operations.
  • Sovereignty: Negligible. The incident did not impact state authority, electoral systems, or government decision-making processes.
  • Economic security: Negligible. While involving a commercial AI startup, the incident did not threaten critical economic infrastructure or strategic technological advantages.
  • Societal stability: Minor. The generation of nonconsensual sexual imagery violates individual privacy and dignity, representing localized societal harm, but is manageable through standard content moderation updates.
Explore in the interactive Incident Tracker