Voice actors were targeted in a harassment campaign where AI-generated voice clones were used to read out their private home addresses on Twitter. While some clips were attributed to ElevenLabs, the company stated that many of the specific clips in this campaign were not created using their software, suggesting a smear campaign or the use of alternative platforms. The incident highlights the risks of AI voice synthesis for individuals with public audio footprints.
Twitter users allegedly used ElevenLab's AI voice synthesis system to impersonate and dox voice actors.
Risk classification
- Primary risk domain: 4 Malicious actors
- Primary risk subdomain: 4.3 Fraud, scams, and targeted manipulation
The incident involves the targeted harassment and humiliation of specific voice actors using synthetic voice clones to read out their private home addresses and offensive slurs.
Additional risk subdomains
- 1.2 Exposure to toxic content: The AI-generated voice clones exposed users and victims to highly toxic content, including hate speech, racist slurs, and homophobic slurs.
Causal factors
- Entity: Human
- Intent: Intentional
- Timing: Post-deployment
The harm was directly caused by the intentional actions of human trolls who used AI voice cloning technology to harass and dox specific victims.
EU AI Act risk tier
- Risk tier: 3 Limited Risk
Limited Risk: The report describes an AI system used to generate deepfakes (voice clones), which falls under Limited Risk and requires specific transparency obligations to ensure users are informed that the content is AI-generated.
AI system and alleged parties
- AI system: ElevenLabs
- AI purpose: Voice Generation; Text Style Replication
- Behaviour type: Tool
- Alleged developer: ElevenLabs
- Alleged deployer: unknown
- Alleged harmed parties: Voice Actors
Harm severity
Highest direct severity in any category: Minor. Severity is scored from Negligible to Catastrophic in each harm category, for harm the reports describe as caused directly or indirectly by the AI system.
- Physical: direct Negligible, indirect Negligible
- Infrastructure: direct Negligible, indirect Negligible
- Property: direct Negligible, indirect Negligible
- Financial: direct Negligible, indirect Negligible
- Environmental: direct Negligible, indirect Negligible
- Malicious content: direct Minor, indirect Minor
- Differential treatment: direct Negligible, indirect Negligible
- Civil rights: direct Negligible, indirect Negligible
- Democracy: direct Negligible, indirect Negligible
- Privacy: direct Minor, indirect Minor
- Psychological: direct Minor, indirect Minor
- Epistemic: direct Negligible, indirect Minor
- Child sexual exploitation and abuse: direct Negligible, indirect Negligible
Malicious content
Reported: The report explicitly describes the creation and spread of highly toxic and offensive audio clips containing racist and homophobic slurs.
Directly caused: AI-generated voice clones were directly used to produce highly toxic audio clips containing racist and homophobic slurs, as well as offensive statements about killing and abusing children.
Indirectly caused: These toxic audio clips were shared and distributed on Twitter, exposing social media users to highly offensive and harmful content.
Inferred additional harm: It is likely that other toxic clips were generated and shared across online forums like 4chan and Twitter, exposing hundreds of users to hate speech and harassment.
Privacy
Reported: The report explicitly describes privacy violations through the publication of private home addresses (doxing).
Directly caused: The AI-generated voices were used to read out the private home addresses of at least four voice actors, violating their privacy.
Indirectly caused: The publication of these addresses on Twitter exposed the victims' private information to the public, increasing their vulnerability to physical harassment.
Inferred additional harm: The ease of cloning voices from public audio footprints poses a significant ongoing threat to the privacy of any individual with online audio recordings.
Psychological
Reported: The report describes targeted harassment and doxing, which inherently causes psychological distress, though it does not explicitly detail clinical diagnoses.
Directly caused: The synthetic voices reading out the victims' home addresses and using highly offensive slurs directly caused distress and alarm to the targeted voice actors.
Indirectly caused: The harassment campaign caused anxiety and safety concerns among the targeted voice actors and the broader voice acting community.
Inferred additional harm: It is highly likely that the four identified victims experienced severe anxiety, fear for their physical safety, and emotional distress due to their private home addresses being publicized alongside hateful rhetoric.
Epistemic
Reported: The report explicitly describes the creation of fake audio clips (voice clones) making it sound like the victims said highly offensive things.
Directly caused: The AI system was used to generate highly realistic fake audio clips of voice actors making offensive statements they never actually said, misrepresenting their identities.
Indirectly caused: The dissemination of these deepfake audio clips on Twitter misled listeners into potentially believing the actors held these offensive views or spoke those words.
Inferred additional harm: The widespread availability of easy-to-use voice cloning technology threatens the overall trust in audio evidence and the authenticity of spoken statements online.
People affected
- Occurrences reported: 1
- People reportedly harmed: 4
- People reportedly exposed: 4
Potential causes
Management
- Gutted Trust and Safety Teams: Twitter's reduced staff led to ineffective content moderation.
- Delayed Safeguard Implementation: ElevenLabs introduced safety features only after abuse occurred.
Technology
- Low Barrier to Voice Cloning: AI tools allowed easy replication of voices with minimal audio samples.
- Lack of AI Platform Safeguards: Initial lack of voice verification allowed unauthorized cloning.
Data Inputs
- Publicly Available Voice Data: Abundant online audio from podcasts and streams used to train models.
- Unprotected Personal Addresses: Doxers easily obtained and input victims' home addresses.
Human Factors
- Malicious Intent of Trolls: Online trolls actively sought to harass and provoke voice actors.
- Public Backlash Against AI: Trolls targeted actors who spoke out against AI voice cloning.
Process and Methods
- Slow Content Moderation: Twitter's support system failed to quickly remove doxing content.
- Inadequate Verification Processes: Lack of validation that user owns the voice being cloned.
Regulatory Environment
- Lack of AI Voice Regulations: No clear legal framework protecting individuals' voice likeness.
- Weak Enforcement of Doxing Laws: Difficulties in legally prosecuting anonymous online harassers.
Information quality
- Classification confidence: High
- Reason for confidence: The report from Motherboard provides detailed first-hand accounts from the victims, direct quotes from the affected voice actors, and statements from the AI developer ElevenLabs. The nature of the harassment, the platforms used, and the specific content of the audio clips are clearly documented.
- Ambiguities identified: There is some ambiguity regarding which specific AI platform was used to generate the doxing clips, as ElevenLabs confirmed their tool was used for the Agent 47 clip but denied being the source of the doxing clips.
- Alternative interpretations: The doxing clips could have been generated using open-source voice cloning tools rather than ElevenLabs' commercial platform.
An online harassment campaign used AI voice cloning to dox and abuse voice actors. While showcasing the evolving capabilities of synthetic media for targeted harassment, the incident remains a localized cybercrime issue with minor national security implications.
- Overall national security impact: Minor
- Response level: Moderate
- Scope: Single nation
- Primary target: USA
- Alleged perpetrator: Unknown
Threat characteristics
- Imminence: Long-term. The specific campaign has concluded, but the underlying capability of AI voice cloning represents an ongoing strategic concern for personal security and fraud.
- Autonomy: Human-controlled. The AI system acted purely as a tool, generating specific audio based on direct text inputs and voice samples provided by human perpetrators.
- Novelty: Evolved capability. Represents an advancement in harassment methods, using highly realistic AI voice synthesis to enhance the psychological impact of traditional doxing.
Impact by dimension
- Physical security: Negligible. The incident involved online harassment and doxing of individuals, presenting no threat to physical systems, critical infrastructure, or national safety.
- Information security: Negligible. While synthetic media was used, this was a localized harassment campaign targeting voice actors rather than a coordinated state-sponsored information warfare operation.
- Sovereignty: Negligible. No impact on government decision-making, electoral systems, state authority, or core government operations.
- Economic security: Negligible. The incident did not target strategic industries, financial systems, or national technological competitiveness.
- Societal stability: Minor. Targeted harassment and doxing of individuals using synthetic voices represents a minor, localized threat to personal privacy and security, manageable under standard cybercrime laws.