AI Tools Failed to Sufficiently Predict COVID Patients, Some Potentially Harmful

During the COVID-19 pandemic, hundreds of AI-based diagnostic and triage tools were developed and deployed in clinical settings. Multiple large-scale studies, including reviews by the Turing Institute and researchers at Maastricht University and the University of Cambridge, concluded that none of these tools were fit for clinical use due to poor data quality, flawed training methodologies, and lack of clinical validation. Experts expressed significant concern that these tools, some of which were marketed to hospitals, may have caused patient harm by providing inaccurate diagnostic information.

AI tools failed to sufficiently predict COVID patients, some potentially harmful.

Source: AI Incident Database

Risk classification

  • Primary risk domain: 7 AI system safety, failures, & limitations
  • Primary risk subdomain: 7.3 Lack of capability or robustness

The AI diagnostic tools lacked the capability and robustness required for clinical use, failing to perform reliably due to flawed training methodologies and poor data quality.

Additional risk subdomains

  • 5.1 Overreliance and unsafe use: Hospitals and clinicians deployed and relied on these unvalidated AI tools in critical clinical situations before they were ready.

Causal factors

  • Entity: AI
  • Intent: Unintentional
  • Timing: Post-deployment

The risk of patient misdiagnosis was caused by the AI systems' incorrect outputs, which was an unintentional outcome of trying to help clinical triage, occurring after these models were deployed or marketed to hospitals.

EU AI Act risk tier

  • Risk tier: 2 High Risk

High Risk: The report describes deep-learning models for diagnosing covid and predicting patient risk from medical images, which falls under AI used in healthcare, such as diagnostic tools.

AI system and alleged parties

  • AI system: predictive tools
  • AI purpose: Medical Diagnosis Support; Image Classification
  • Behaviour type: Assistant
  • Alleged developer: unknown
  • Alleged deployer: unknown
  • Alleged harmed parties: Doctors, COVID patients

Harm severity

Highest direct severity in any category: Substantial. Severity is scored from Negligible to Catastrophic in each harm category, for harm the reports describe as caused directly or indirectly by the AI system.

  • Physical: direct Negligible, indirect Negligible
  • Infrastructure: direct Negligible, indirect Negligible
  • Property: direct Negligible, indirect Negligible
  • Financial: direct Negligible, indirect Negligible
  • Environmental: direct Negligible, indirect Negligible
  • Malicious content: direct Negligible, indirect Negligible
  • Differential treatment: direct Negligible, indirect Negligible
  • Civil rights: direct Negligible, indirect Negligible
  • Democracy: direct Negligible, indirect Negligible
  • Privacy: direct Negligible, indirect Negligible
  • Psychological: direct Negligible, indirect Negligible
  • Epistemic: direct Minor, indirect Negligible
  • Child sexual exploitation and abuse: direct Negligible, indirect Negligible

Epistemic

Reported: The report describes how flawed AI models generated misleading diagnostic and risk predictions, though this was in a clinical rather than public information context.

Directly caused: N/A

Indirectly caused: N/A

Inferred additional harm: The flood of hundreds of mediocre, non-functioning AI models may have contributed to clinical skepticism and eroded trust in medical AI technologies.

People affected

  • Occurrences reported: 1
  • People reportedly exposed: 1000

Potential causes

Management

  • Secrecy and NDAs: Hospitals signed NDAs with vendors, preventing independent model evaluation.
  • Poor Academic Incentives: Lack of career incentives to share work or validate existing clinical models.

Technology

  • Flawed Model Training: Models learned patient position or text fonts instead of actual clinical signs.
  • Overfitting on Duplicated Data: Frankenstein datasets caused models to be tested on their own training data.

Data Inputs

  • Poor Quality Frankenstein Datasets: Datasets spliced from multiple sources contained duplicates and unknown origins.
  • Biased Training Images: Healthy control scans were children, teaching AI to detect age, not COVID.
  • Incorporation Bias in Labels: Scans were labeled by doctor opinion rather than objective PCR tests.

Human Factors

  • Lack of Cross-Disciplinary Skills: AI developers lacked clinical insight and medical experts lacked math skills.
  • Hype and Unrealistic Expectations: High hype encouraged deployment of unready tools, risking patient safety.

Process and Methods

  • Lack of Model Validation: Researchers rushed to build new models instead of validating existing ones.
  • Non-Standardized Data Formats: Absence of standardized data formats hindered effective data sharing.

Regulatory Environment

  • Lack of Clinical Oversight: Tools were marketed and used without rigorous clinical validation or standards.

Information quality

  • Classification confidence: High
  • Reason for confidence: The reports provide high-quality, peer-reviewed evidence from major studies (BMJ, Nature Machine Intelligence) detailing the exact failure modes of the AI models. While the exact number of patients affected is not quantified, the technical and systemic aspects of the failure are exceptionally well-documented.
  • Ambiguities identified: The exact number of hospitals using these tools and the precise number of patients who may have been harmed remains undisclosed due to nondisclosure agreements and secrecy.

During the COVID-19 pandemic, hundreds of unvalidated AI diagnostic and triage tools were deployed in clinical settings globally. Flawed training methodologies and data quality issues rendered them unfit for clinical use, posing minor national security risks by potentially compromising critical healthcare infrastructure and patient safety during a public health crisis.

  • Overall national security impact: Minor
  • Response level: Moderate
  • Scope: Multiple nations
  • Primary target: No clear primary
  • Other affected: European nations
  • Alleged perpetrator: Unknown

Threat characteristics

  • Imminence: Long-term. The incident represents a long-term strategic and regulatory concern regarding the validation and safety of AI in critical sectors like healthcare.
  • Autonomy: Human-supervised. The AI tools were designed to assist clinicians in triage and diagnosis, operating with human oversight before clinical decisions were finalized.
  • Novelty: Evolved capability. While clinical decision support failures have occurred, the rapid, global-scale deployment of hundreds of unvalidated models during a pandemic was an evolved threat.

Impact by dimension

  • Physical security: Minor. Widespread deployment of unvalidated AI diagnostic tools in hospitals during a pandemic threatened critical healthcare infrastructure, with potential to cause patient harm through misdiagnosis.
  • Information security: Negligible. No evidence of information warfare, intelligence compromise, or hostile disinformation campaigns associated with this incident.
  • Sovereignty: Negligible. Core government operations and sovereignty were not threatened, as the issue was restricted to clinical decision support tool failures.
  • Economic security: Negligible. While research resources were wasted on redundant models, there was no significant threat to national economic stability or strategic technology theft.
  • Societal stability: Negligible. No systemic human rights violations, mass surveillance, or widespread civil unrest resulted from the deployment of these tools.
Explore in the interactive Incident Tracker