Facebook Messenger AI Stickers Generate Ethical and Content Moderation Concerns

Meta's AI-generated sticker tool, powered by Llama 2, was found to produce inappropriate and offensive content, including sexualized imagery and depictions of public figures in compromising situations. Users discovered that the tool's safety filters could be bypassed using typos or specific prompts, leading to the generation of problematic stickers. This incident highlights challenges in content moderation for generative AI features deployed in consumer messaging platforms.

Facebook Messenger AI stickers, a feature by Meta, allows users to generate personalized stickers via AI for use in conversations. While the feature has been praised for its creativity, it has also stirred controversy for its alleged production of inappropriate or offensive content. This has raised questions about the effectiveness of Meta's content moderation measures and the ethical responsibilities associated with AI-driven content generation.

Source: AI Incident Database

Risk classification

  • Primary risk domain: 1 Discrimination & Toxicity
  • Primary risk subdomain: 1.2 Exposure to toxic content

The AI tool generated and exposed users to highly inappropriate, offensive, and sexualized content, including child soldiers and nude illustrations of public figures.

Additional risk subdomains

  • 7.3 Lack of capability or robustness: The system's safety filters were easily bypassed using typos or descriptive workarounds, showing a lack of robustness in content moderation.

Causal factors

  • Entity: AI
  • Intent: Unintentional
  • Timing: Post-deployment

The generation of inappropriate and offensive stickers was an unexpected and unintentional outcome of deploying the Llama 2-powered sticker tool.

EU AI Act risk tier

  • Risk tier: 3 Limited Risk

Limited Risk: The system is an AI-generated content tool used to create synthetic images (stickers), which falls under the category of AI-generated content requiring transparency obligations.

AI system and alleged parties

  • AI system: AI-generated sticker tool 2 (Meta)
  • AI purpose: Image Generation; Social Media Content Generation
  • Behaviour type: Tool
  • Alleged developer: Meta
  • Alleged deployer: Meta
  • Alleged harmed parties: Facebook Messenger users

Harm severity

Highest direct severity in any category: Minor. Severity is scored from Negligible to Catastrophic in each harm category, for harm the reports describe as caused directly or indirectly by the AI system.

  • Physical: direct Negligible, indirect Negligible
  • Infrastructure: direct Negligible, indirect Negligible
  • Property: direct Negligible, indirect Negligible
  • Financial: direct Negligible, indirect Negligible
  • Environmental: direct Negligible, indirect Negligible
  • Malicious content: direct Minor, indirect Negligible
  • Differential treatment: direct Negligible, indirect Negligible
  • Civil rights: direct Negligible, indirect Negligible
  • Democracy: direct Negligible, indirect Negligible
  • Privacy: direct Negligible, indirect Negligible
  • Psychological: direct Negligible, indirect Negligible
  • Epistemic: direct Negligible, indirect Negligible
  • Child sexual exploitation and abuse: direct Negligible, indirect Negligible

Malicious content

Reported: The report explicitly describes the generation of inappropriate and offensive sticker images, including child soldiers, nude illustrations of public figures, and sexualized cartoon characters.

Directly caused: The Llama 2-powered sticker tool directly generated offensive images such as child soldiers, nude illustrations of Canadian Prime Minister [PERSON_002], and sexualized depictions of Sonic the Hedgehog.

Indirectly caused: N/A

Inferred additional harm: It is likely that many more inappropriate or offensive stickers were generated and shared among the select group of users who had access to the tool during its early rollout.

People affected

  • Occurrences reported: 1
  • People reportedly exposed: 5

Potential causes

Management

  • Risky Feature Rollout: Meta deployed the tool to select users despite unresolved safety issues.

Technology

  • Weak Algorithmic Safeguards: The model generated inappropriate images like nude figures and weapons.
  • Easily Bypassed Filters: Users bypassed blocked words using simple typos or descriptive text.

Data Inputs

  • Unfiltered Prompt Inputs: Certain prompts like 'World Trade Center' bypassed filters without warnings.

Human Factors

  • Adversarial User Testing: Users intentionally entered prompts to test limits and generate bad images.

Process and Methods

  • Insufficient Pre-Release Testing: Live user testing was used to find obvious vulnerabilities in the tool.

Information quality

  • Classification confidence: High
  • Reason for confidence: The report clearly details the AI system involved (Meta's Llama 2-powered sticker tool), the platform (Facebook Messenger), the nature of the generated content, and how the safety filters were bypassed. There is little ambiguity regarding the events.
  • Ambiguities identified: The exact number of users who had access to the tool during this early rollout is not specified.
  • Alternative interpretations: None. The incident is a straightforward case of generative AI safety filter failure leading to inappropriate content generation.

Meta's Llama 2-powered sticker tool generated highly inappropriate and offensive stickers, including depictions of child soldiers and nude illustrations of Canadian Prime Minister Justin Trudeau, due to easily bypassed safety filters. The national security impact remains minor, primarily highlighting the ongoing challenges of content moderation and information security in consumer-facing generative AI tools.

  • Overall national security impact: Minor
  • Response level: Moderate
  • Scope: Multiple nations
  • Primary target: No clear primary
  • Other affected: Canada, United States
  • Alleged perpetrator: Unknown

Threat characteristics

  • Imminence: Long-term. This represents an ongoing, long-term challenge of content moderation and safety filters in consumer-facing generative AI rather than an immediate national security crisis.
  • Autonomy: Human-controlled. The AI functioned strictly as a tool, generating images directly in response to user text prompts and deliberate bypass attempts.
  • Novelty: Evolved capability. The incident demonstrates an evolved capability of users finding novel workarounds (like typos and descriptive prompts) to bypass standard safety filters in newly deployed text-to-image models.

Impact by dimension

  • Physical security: Negligible. The incident involved a sticker-generation tool and posed no threat to physical systems, infrastructure, or human safety.
  • Information security: Minor. The tool generated inappropriate illustrations of public figures like Canada's Prime Minister, representing a minor information security concern, though it lacked the coordination of a state-sponsored disinformation campaign.
  • Sovereignty: Negligible. No disruption to state authority, electoral systems, or core government decision-making processes was reported.
  • Economic security: Negligible. The incident did not involve financial attacks, critical supply chain vulnerabilities, or the theft of strategic technology.
  • Societal stability: Minor. The creation of offensive stickers, including depictions of child soldiers and sexualized characters, represents a minor societal concern but did not cause large-scale civil unrest or systemic human rights violations.
Explore in the interactive Incident Tracker