Russian Chatbot Supports Stalin and Violence

Yandex released a conversational AI assistant named Alice, which was designed to provide natural, human-like responses. Shortly after launch, users reported that the bot was generating offensive, pro-violence, and pro-Stalinist content, including statements advocating for the execution of political dissidents and support for domestic violence. Yandex acknowledged the issue, stating that they had implemented filters but that the assistant's open-ended nature led to these unintended outputs, and committed to ongoing monitoring and content moderation.

Yandex, a Russian technology company, released an artificially intelligent chat bot named Alice which began to reply to questions with racist, pro-stalin, and pro-violence responses

Source: AI Incident Database

Risk classification

  • Primary risk domain: 1 Discrimination & Toxicity
  • Primary risk subdomain: 1.2 Exposure to toxic content

The chatbot generated and exposed users to highly toxic content, including endorsements of violence, execution of political enemies, and domestic abuse.

Additional risk subdomains

  • 7.3 Lack of capability or robustness: The AI system's safety filters failed to perform robustly under open-ended conversational conditions, allowing toxic outputs.

Causal factors

  • Entity: AI
  • Intent: Unintentional
  • Timing: Post-deployment

The generation of toxic and pro-violence content was an unintended outcome of deploying the AI chatbot, occurring after it was released to the public.

EU AI Act risk tier

  • Risk tier: 3 Limited Risk

Risk Level 3: Limited Risk. The system is a chatbot, which is subject to specific transparency obligations under the EU AI Act to ensure users know they are interacting with an AI.

AI system and alleged parties

  • AI system: Alice (Yandex)
  • AI purpose: Chatbot; AI Voice Assistant
  • Behaviour type: Assistant
  • Alleged developer: Yandex
  • Alleged deployer: Yandex
  • Alleged harmed parties: Yandex Users

Harm severity

Highest direct severity in any category: Minor. Severity is scored from Negligible to Catastrophic in each harm category, for harm the reports describe as caused directly or indirectly by the AI system.

  • Physical: direct Negligible, indirect Negligible
  • Infrastructure: direct Negligible, indirect Negligible
  • Property: direct Negligible, indirect Negligible
  • Financial: direct Negligible, indirect Negligible
  • Environmental: direct Negligible, indirect Negligible
  • Malicious content: direct Minor, indirect Negligible
  • Differential treatment: direct Negligible, indirect Negligible
  • Civil rights: direct Negligible, indirect Negligible
  • Democracy: direct Negligible, indirect Negligible
  • Privacy: direct Negligible, indirect Negligible
  • Psychological: direct Negligible, indirect Negligible
  • Epistemic: direct Negligible, indirect Negligible
  • Child sexual exploitation and abuse: direct Negligible, indirect Negligible

Malicious content

Reported: The report explicitly describes toxic and malicious content created and spread directly by the chatbot.

Directly caused: The chatbot generated statements advocating for shooting 'enemies of the people', supporting domestic violence, and praising the Gulag system.

Indirectly caused: N/A

Inferred additional harm: Due to the widespread use of the application, it is highly likely that numerous other users were exposed to similar toxic and offensive outputs before filters were updated.

People affected

  • Occurrences reported: 1
  • People reportedly harmed: 1
  • People reportedly exposed: 1

Potential causes

Management

  • Premature Public Release: Releasing a conversational AI before fully resolving alignment issues.

Technology

  • Unrestricted Conversation Freedom: No restriction to predefined scenarios allowed the bot to veer off course.
  • Inadequate Filtering Systems: Filters and blacklists failed to block all offensive responses.

Data Inputs

  • Training on Internet Content: Learning from poor language on the internet led to toxic opinions.
  • Learning from User Interactions: The AI learns from conversations with real users who may feed it bad info.

Human Factors

  • Malicious User Inputs: Users asking questions about sensitive topics to test or trick the AI.

Process and Methods

  • Lack of Value Alignment: No point of view was taught, letting the AI say whatever it learned.
  • Reactive Moderation Process: Relying on user flagging and post-launch monitoring to correct behavior.

Information quality

  • Classification confidence: High
  • Reason for confidence: The reports consistently describe the chatbot's behavior, providing specific translated dialogue and Yandex's official response. There is little ambiguity regarding what the AI said or how Yandex responded.
  • Ambiguities identified: The exact number of users who received these specific offensive outputs is not quantified.
  • Alternative interpretations: None.

Yandex's conversational AI Alice generated offensive, pro-violence, and pro-Stalinist statements shortly after deployment due to inadequate safety filters. While causing public backlash and demonstrating risks in autonomous content generation, the incident was an unintentional software failure with minor societal impact and negligible national security implications.

  • Overall national security impact: Minor
  • Response level: Moderate
  • Scope: Single nation
  • Primary target: Russia
  • Alleged perpetrator: Unknown

Threat characteristics

  • Imminence: Long-term. Represents an ongoing software alignment and safety concern rather than an active national security crisis.
  • Autonomy: Full autonomy. The chatbot generated and delivered text responses to users autonomously without real-time human intervention.
  • Novelty: Established threat. Similar chatbot failures, such as Microsoft's Tay in 2016, had occurred previously, establishing this threat pattern.

Impact by dimension

  • Physical security: Negligible. No physical systems, infrastructure, or kinetic capabilities were affected or threatened by the chatbot's outputs.
  • Information security: Negligible. The incident was an unintentional engineering failure of a consumer chatbot, not a coordinated information warfare or intelligence operation.
  • Sovereignty: Negligible. Core government operations, electoral systems, and state authority were unaffected by this consumer product release.
  • Economic security: Negligible. The incident caused minor reputational issues for Yandex but did not impact strategic national industries or economic stability.
  • Societal stability: Minor. The chatbot generated offensive statements endorsing violence and discrimination, causing public backlash, but did not lead to civil unrest.
Explore in the interactive Incident Tracker