False Negatives for Water Quality-Associated Beach Closures

Toronto Public Health contracted with Cann Forecast to use an AI predictive model to monitor water quality at two beaches. The model consistently underperformed, failing to identify dangerous E. coli levels and resulting in beaches remaining open when they should have been closed. This failure exposed beachgoers and staff to potential health risks, highlighting issues with over-reliance on automated risk prediction tools without adequate human oversight or validation.

Toronto’s use of AI predictive modeling (AIPM) which had replaced existing methodology as the only determiner of beach water quality raised concerns about its accuracy, after allegedly conflicting results were found by a local water advocacy group using traditional means.

Source: AI Incident Database

Risk classification

  • Primary risk domain: 7 AI system safety, failures, & limitations
  • Primary risk subdomain: 7.3 Lack of capability or robustness

The AI model failed to perform reliably, predicting safe water quality when E. coli levels were dangerously high, representing a clear lack of capability and robustness.

Additional risk subdomains

  • 5.1 Overreliance and unsafe use: City officials exhibited automation bias, allowing the unproven AI model's predictions to supersede actual daily laboratory testing results.

Causal factors

  • Entity: AI
  • Intent: Unintentional
  • Timing: Post-deployment

The risk was caused by the inaccurate predictions of the deployed Cann Forecast AI model, which failed to perform as expected after being put into active use.

EU AI Act risk tier

  • Risk tier: 2 High Risk

High Risk: The AI system is used in public safety and environmental monitoring (water quality at public beaches), which has significant implications for public health and safety.

AI system and alleged parties

  • AI system: Cann Forecast predictive modeling (Cann Forecast)
  • AI purpose: Substance Detection; Value Estimation
  • Behaviour type: Tool
  • Alleged developer: Toronto Public Health
  • Alleged deployer: Toronto city government
  • Alleged harmed parties: Toronto citizens, Sunnyside beachgoers, Marie Curtis beachgoers

Harm severity

Highest direct severity in any category: Minor. Severity is scored from Negligible to Catastrophic in each harm category, for harm the reports describe as caused directly or indirectly by the AI system.

  • Physical: direct Negligible, indirect Negligible
  • Infrastructure: direct Negligible, indirect Negligible
  • Property: direct Negligible, indirect Negligible
  • Financial: direct Minor, indirect Negligible
  • Environmental: direct Negligible, indirect Negligible
  • Malicious content: direct Negligible, indirect Negligible
  • Differential treatment: direct Negligible, indirect Negligible
  • Civil rights: direct Negligible, indirect Negligible
  • Democracy: direct Negligible, indirect Negligible
  • Privacy: direct Negligible, indirect Negligible
  • Psychological: direct Negligible, indirect Negligible
  • Epistemic: direct Negligible, indirect Negligible
  • Child sexual exploitation and abuse: direct Negligible, indirect Negligible

Financial

Reported: The report describes the CA$30,000 cost of the contract awarded to Cann Forecast for the pilot program.

Directly caused: The city of Toronto spent CA$30,000 on the failed AI predictive modeling pilot program.

Indirectly caused: N/A

Inferred additional harm: N/A

People affected

  • Occurrences reported: 1
  • People reportedly exposed: 1000

Potential causes

Management

  • Prioritizing Automation: TPH adopted AI to replace lagging lab tests without adequate risk assessment.
  • Deflecting Public Concerns: Management falsely claimed human oversight to deflect from model failures.
  • Lack of Operational Transparency: The city quietly adopted the tool without public disclosure or metrics.

Technology

  • Faulty Weather Data: The model used incorrect weather data to make predictions at Sunnyside Beach.
  • Inherent Model Inaccuracy: The algorithm was less accurate than a coin flip at identifying unsafe days.
  • Data Leakage: Leakage in similar tools led to exaggerated initial claims of performance.

Data Inputs

  • Incorrect Weather Inputs: Sunnyside predictions were compromised by ingestion of faulty weather data.
  • Lagging Validation Data: Decisions were made before validating real-time bacteria levels.
  • Biased Training Data: In related tools, biased data inputs led to systematic discrimination.

Human Factors

  • Automation Bias: Staff defaulted to the automated tool, failing to challenge faulty outputs.
  • Lack of User Expertise: Non-expert officials bought the tool on faith without ability to evaluate.
  • Moral Crumple Zones: Blaming human operators for failures of systems they cannot safely oversee.

Process and Methods

  • No Performance Tracking: City officials never checked the deployed model's performance over summer.
  • Superseding Reliable Tests: AI predictions superseded reliable, traditional laboratory E. coli testing.
  • Lack of Human Override: Staff never overrode the model despite claims of human-in-the-loop process.

Regulatory Environment

  • Noncompetitive Procurement: The city awarded a CA$30,000 contract without any competing bids.
  • Proprietary Trade Secrets: Vendors shield models from external scrutiny under trade secret claims.
  • Lack of AI Safety Standards: No standard regulatory framework existed to validate high-stakes AI tools.

Information quality

  • Classification confidence: High
  • Reason for confidence: The reports provide clear, detailed, and consistent accounts of the Toronto beach water quality AI pilot, including specific statistics on accuracy, dates, and the nature of the failure. There is high agreement between the sources on the core facts of the incident.
  • Ambiguities identified: None of significance; the role of the AI and the nature of its failure are explicitly documented.
  • Alternative interpretations: None; the failure of the predictive model is clearly established by both independent analysis and admission by public health officials.

A municipal-level failure of an AI water quality predictive pilot in Toronto, Canada, resulted in localized public health risks due to E. coli exposure. The incident highlights issues of automation bias and poor data validation in local government procurement, but carries negligible national security implications.

  • Overall national security impact: Negligible
  • Response level: Minor
  • Scope: Single nation
  • Primary target: Canada
  • Alleged perpetrator: Unknown

Threat characteristics

  • Imminence: Long-term. The incident represents a past pilot program failure and poses no immediate or near-term national security threat.
  • Autonomy: Human-supervised. The AI operated as a predictive tool with humans retaining ultimate authority to make beach closure decisions, despite exhibiting automation bias.
  • Novelty: Established threat. Failures due to inaccurate models, faulty data inputs, and automation bias are well-documented issues in predictive AI applications.

Impact by dimension

  • Physical security: Negligible. The failure of a water quality prediction model at municipal beaches is a local public health issue, posing no threat to critical national infrastructure or physical security.
  • Information security: Negligible. No elements of information warfare, espionage, or intelligence compromise are present in this municipal pilot program.
  • Sovereignty: Negligible. The incident only affected local municipal decision-making regarding beach closures, with no impact on national sovereignty or federal government functions.
  • Economic security: Negligible. The financial cost of 30,000 Canadian dollars is negligible and poses no threat to national economic stability or strategic technological competitive advantage.
  • Societal stability: Negligible. Localized public health risks and anxiety among lifeguards do not constitute a population-scale threat to societal stability or civil liberties.
Explore in the interactive Incident Tracker