European researchers conducted a study on Microsoft Copilot's medical advice, finding that only 54% of responses were scientifically accurate. The study highlighted that a significant portion of the AI's answers could lead to moderate or severe harm, including risks of death. The report serves as a warning against relying on AI search tools for critical medical information.
Microsoft Copilot, when asked medical questions, was reportedly found to provide accurate information only 54% of the time, according to European researchers (citation provided in editor's notes). Analysis by the researchers reported that 42% of Copilot's responses could cause moderate to severe harm, with 22% of responses posing a risk of death or severe injury.
Risk classification
- Primary risk domain: 7 AI system safety, failures, & limitations
- Primary risk subdomain: 7.3 Lack of capability or robustness
The incident involves Microsoft Copilot failing to provide accurate medical information in a critical application, demonstrating a lack of capability and robustness that could lead to severe patient harm.
Additional risk subdomains
- 3.1 False or misleading information: The AI system generated scientifically inaccurate and wrong medical answers, spreading false information to users.
- 5.1 Overreliance and unsafe use: Users may overrely on AI search tools for critical medical advice due to accessibility or affordability issues, despite the systems being unfit for this purpose.
Causal factors
- Entity: AI
- Intent: Unintentional
- Timing: Post-deployment
The risk arises from inaccurate medical advice generated by the deployed Microsoft Copilot AI system, which is an unintentional failure of the model's output generation.
EU AI Act risk tier
- Risk tier: 3 Limited Risk
Limited Risk: The systems described are general-purpose chatbots and AI search assistants, which fall under Risk Level 3 and carry transparency obligations to ensure users know they are interacting with AI.
AI system and alleged parties
- AI system: Google AI search, Microsoft Copilot (Google, Microsoft)
- AI purpose: Question Answering; Content Search
- Behaviour type: Assistant
- Alleged developer: Microsoft
- Alleged deployer: Microsoft Copilot, Microsoft
- Alleged harmed parties: People seeking medical advice, Microsoft Copilot users, General public
Harm severity
Highest direct severity in any category: Substantial. Severity is scored from Negligible to Catastrophic in each harm category, for harm the reports describe as caused directly or indirectly by the AI system.
- Physical: direct Negligible, indirect Negligible
- Infrastructure: direct Negligible, indirect Negligible
- Property: direct Negligible, indirect Negligible
- Financial: direct Negligible, indirect Negligible
- Environmental: direct Negligible, indirect Negligible
- Malicious content: direct Negligible, indirect Negligible
- Differential treatment: direct Negligible, indirect Negligible
- Civil rights: direct Negligible, indirect Negligible
- Democracy: direct Negligible, indirect Negligible
- Privacy: direct Minor, indirect Negligible
- Psychological: direct Negligible, indirect Negligible
- Epistemic: direct Minor, indirect Minor
- Child sexual exploitation and abuse: direct Negligible, indirect Negligible
Privacy
Reported: The report explicitly describes a privacy violation where Google's AI search listed a private citizen's phone number.
Directly caused: Google's AI search arbitrarily listed a private citizen's phone number as the corporate headquarters phone number for a video game publisher.
Indirectly caused: N/A
Inferred additional harm: It is possible that other private individuals have had their personal contact information or sensitive data mistakenly exposed by AI search summaries indexing public or private web data.
Epistemic
Reported: The report explicitly describes epistemic harm in the form of AI-driven misinformation and inaccurate medical advice.
Directly caused: Microsoft Copilot generated scientifically inaccurate medical information in 46% of tested cases, with 3% of answers being completely wrong and 24% not matching established medical knowledge.
Indirectly caused: Google AI search generated misleading recommendations such as telling users to eat rocks or add glue to pizza, and hallucinated that there are 150 Planet Hollywood restaurants in Guam.
Inferred additional harm: It is highly likely that many users seeking medical or general information online have been misled by these inaccurate AI search summaries, potentially eroding trust in online information.
People affected
- Occurrences reported: 1
- People reportedly exposed: 1
Potential causes
Management
- Premature Deployment of AI: Releasing AI search features for medical queries despite known errors.
Technology
- Inaccurate Model Outputs: Copilot generated inaccurate medical advice 46 percent of the time.
- Severe Generation Errors: Three percent of the generated medical answers were completely wrong.
- High Potential for Harm: Twenty-two percent of answers could lead to death or severe harm.
Data Inputs
- Unaligned Medical Knowledge: AI answers failed to match established medical knowledge in 24% of cases.
Human Factors
- User Reliance on AI Search: Users consult AI as first point-of-call due to costly healthcare access.
- Taking AI Answers at Face Value: Users may trust and act on harmful medical advice without verification.
Process and Methods
- Ineffective Disclaimers: Fine print telling users to check answers shifts safety responsibility.
Regulatory Environment
- Lack of Regulatory Enforcement: Absence of strict oversight for AI-generated medical information.
Information quality
- Classification confidence: High
- Reason for confidence: The report clearly details the findings of a scientific study on Microsoft Copilot's medical accuracy, providing specific percentages and metrics. It also provides concrete examples of Google AI search failures, making the nature of the AI safety risks highly transparent and easy to classify.
- Ambiguities identified: The report does not specify the exact methodology of the European study beyond the sample size of 500 answers, nor does it provide details on whether any real-world patients have actually suffered harm from acting on Copilot's advice.
- Alternative interpretations: The incident could be interpreted as primarily an overreliance issue (Domain 5.1) rather than a system capability failure (Domain 7.3), but the high rate of scientifically inaccurate outputs supports classifying it primarily as a lack of robustness.
A scientific study revealed that Microsoft Copilot provides accurate medical advice only 54 percent of the time, with potential for severe harm or death in extreme cases. While presenting significant public health and misinformation concerns, the incident has negligible direct national security implications as it lacks state-sponsored intent, kinetic capabilities, or critical infrastructure impact.
- Overall national security impact: Minor
- Response level: Moderate
- Scope: Multiple nations
- Primary target: No clear primary
- Other affected: Europe and United States
- Alleged perpetrator: Unknown
Threat characteristics
- Imminence: Long-term. Represents an ongoing strategic concern regarding the reliability of commercial AI models rather than an active, imminent national security crisis.
- Autonomy: Human-controlled. The AI functions as an informational assistant; humans make the final decisions on whether to act on the generated medical advice.
- Novelty: Established threat. AI hallucinations and search engine inaccuracies are well-known and documented phenomena, representing an established threat rather than a novel capability.
Impact by dimension
- Physical security: Negligible. The report highlights potential patient harm from inaccurate medical advice, but there is no direct threat to critical infrastructure, physical security systems, or state-directed kinetic attacks.
- Information security: Negligible. While the AI generates inaccurate information, this is an unintentional post-deployment capability limitation rather than a coordinated information warfare or intelligence compromise campaign.
- Sovereignty: Negligible. The incident involves commercial search tools and does not impact state authority, electoral processes, or core government decision-making.
- Economic security: Negligible. While reflecting intense market competition between major tech firms, the incident does not involve strategic technology theft, financial system attacks, or critical supply chain disruptions.
- Societal stability: Minor. Widespread reliance on inaccurate AI medical advice could lead to public health concerns, particularly for vulnerable populations lacking healthcare access, though it does not constitute systematic oppression.