A user reported that PayPal's generative AI chatbot falsely informed them of a declined transaction that never occurred. The user was forced to contact human customer support to verify the account status, confirming the chatbot had fabricated the alert. This incident highlights issues with deploying generative AI in financial services without adequate fact-checking or error-handling mechanisms.
On July 13th, 2024, a user reported an incident involving PayPal's generative AI chatbot. The chatbot allegedly incorrectly informed the user of a declined transaction that never occurred, causing confusion and prompting a call to customer service for clarification. This false alert suggests a flaw in the AI system's reliability. The incident created unnecessary labor for both the user and PayPal's human support, demonstrating the potential harm of deploying generative AI without thorough testing and error handling mechanisms.
Risk classification
- Primary risk domain: 3 Misinformation
- Primary risk subdomain: 3.1 False or misleading information
The chatbot fabricated a declined transaction that never occurred, presenting false information to the user.
Additional risk subdomains
- 7.3 Lack of capability or robustness: The chatbot failed to perform reliably by hallucinating data and failing to access actual transaction records.
Causal factors
- Entity: AI
- Intent: Unintentional
- Timing: Post-deployment
The incident was caused by the PayPal Assistant chatbot post-deployment, which unintentionally fabricated a declined transaction alert during a user interaction.
EU AI Act risk tier
- Risk tier: 3 Limited Risk
Limited Risk: The system is a chatbot ('PayPal Assistant'), which falls under Limited Risk due to transparency obligations to ensure users are aware they are interacting with an AI.
AI system and alleged parties
- AI system: PayPal Assistant
- AI purpose: Chatbot; Question Answering
- Behaviour type: Assistant
- Alleged developer: PayPal
- Alleged deployer: PayPal
- Alleged harmed parties: PayPal customers, PayPal customer service representatives, Kiri Wagstaff
Harm severity
Highest direct severity in any category: Minor. Severity is scored from Negligible to Catastrophic in each harm category, for harm the reports describe as caused directly or indirectly by the AI system.
- Physical: direct Negligible, indirect Negligible
- Infrastructure: direct Negligible, indirect Negligible
- Property: direct Negligible, indirect Negligible
- Financial: direct Negligible, indirect Negligible
- Environmental: direct Negligible, indirect Negligible
- Malicious content: direct Negligible, indirect Negligible
- Differential treatment: direct Negligible, indirect Negligible
- Civil rights: direct Negligible, indirect Negligible
- Democracy: direct Negligible, indirect Negligible
- Privacy: direct Negligible, indirect Negligible
- Psychological: direct Negligible, indirect Negligible
- Epistemic: direct Negligible, indirect Negligible
- Child sexual exploitation and abuse: direct Negligible, indirect Negligible
People affected
- Occurrences reported: 1
- People reportedly harmed: 1
- People reportedly exposed: 1
Potential causes
Management
- Premature Deployment Decision: Management deployed the beta chatbot to customers despite known risks.
- Poor Customer Support Loop: Support loop redirected users back to the faulty chatbot.
Technology
- Generative AI Hallucination: The chatbot fabricated a non-existent declined transaction of 23.64 USD.
- Inherent LLM Limitations: Generative AI employs randomness and is not designed to look up facts.
Data Inputs
- Lack of Verifiable Data Sync: The chatbot did not verify actual database records before claiming a decline.
Human Factors
- Deployer Misunderstanding: Deploying AI technology without understanding its core limitations.
Process and Methods
- Lack of Product Testing: The AI product was deployed to customers without sufficient testing.
- Inactive Feedback Channels: The provided support email was inactive, preventing error reporting.
Information quality
- Classification confidence: High
- Reason for confidence: The report is a first-hand account of a specific interaction with the PayPal chatbot, clearly detailing the fabricated transaction and the subsequent verification with human customer support. The details are straightforward and unambiguous.
A PayPal customer service chatbot hallucinated a declined transaction alert, causing minor inconvenience to a user. This incident represents a standard commercial software malfunction with negligible national security implications.
- Overall national security impact: Negligible
- Response level: Minor
- Scope: Unknown
- Primary target: Unknown
- Alleged perpetrator: Unknown
Threat characteristics
- Imminence: Long-term. This is a routine post-deployment software issue rather than an active national security crisis.
- Autonomy: Full autonomy. The chatbot generated and delivered the fabricated transaction alert autonomously without human review prior to sending.
- Novelty: Established threat. Generative AI hallucinations and errors in customer service chatbots are well-documented and common occurrences.
Impact by dimension
- Physical security: Negligible. The incident involved a customer service chatbot hallucination and posed no threat to physical systems, critical infrastructure, or human safety.
- Information security: Negligible. There is no evidence of coordinated information warfare, intelligence compromise, or systematic disinformation targeting national security.
- Sovereignty: Negligible. The incident was confined to a commercial customer service interaction and did not impact state authority, elections, or core government operations.
- Economic security: Negligible. While causing minor inconvenience and wasted customer service time, the incident poses no threat to national economic stability or strategic industries.
- Societal stability: Negligible. The incident was a localized software error affecting a single reported user, with no implications for societal stability or civil liberties.