During the October 2020 California bar exam, an AI-powered proctoring system flagged over one-third of the 9,301 test-takers for alleged rule violations. This high rate of flagging led to widespread accusations of cheating, forcing many applicants to undergo a stressful review process to prove their innocence. Critics and legal representatives argued that the system was overly sensitive and that the process for challenging these allegations was opaque and unfair, as applicants were often denied access to the video evidence used against them.
The proctoring algorithm used in a California bar exam cited a third of thousands of applicants as cheaters, resulting in allegations where exam takers were instructed to prove otherwise without seeing their incriminating video evidence.
Risk classification
- Primary risk domain: 7 AI system safety, failures, & limitations
- Primary risk subdomain: 7.3 Lack of capability or robustness
The AI proctoring software failed to perform reliably under normal test-taking conditions, resulting in an extremely high false-positive rate where 34% of innocent applicants were flagged for cheating.
Additional risk subdomains
- 1.3 Unequal performance across groups: The reports note that the algorithm flagged Black and other minority applicants for merely existing, indicating potential demographic bias and unequal performance.
- 5.1 Overreliance and unsafe use: The bar examiners over-relied on the AI's automated flags, issuing violation notices to thousands of applicants without first conducting human verification of the video evidence.
Causal factors
- Entity: AI
- Intent: Unintentional
- Timing: Post-deployment
The incident was caused by the ExamSoft AI proctoring system unintentionally flagging a massive number of innocent test-takers due to over-sensitivity, occurring after the software was deployed for the online bar exam.
EU AI Act risk tier
High Risk: The AI system is used for professional licensing and vocational assessment, which directly affects access to employment. Under the EU AI Act, AI systems used in education, vocational training, or employment management are classified as High Risk.
AI system and alleged parties
- AI system: ExamSoft proctoring software (ExamSoft)
- AI purpose: Cheating Detection; Workforce Monitoring and Evaluation
- Behaviour type: Autonomous
- Alleged developer: ExamSoft
- Alleged deployer: California Bar’s Committee of Bar Examiners
- Alleged harmed parties: flagged California bar exam takers, California bar exam takers
Harm severity
Highest direct severity in any category: Substantial. Severity is scored from Negligible to Catastrophic in each harm category, for harm the reports describe as caused directly or indirectly by the AI system.
- Physical: direct Negligible, indirect Negligible
- Infrastructure: direct Negligible, indirect Negligible
- Property: direct Negligible, indirect Negligible
- Financial: direct Negligible, indirect Minor
- Environmental: direct Negligible, indirect Negligible
- Malicious content: direct Negligible, indirect Negligible
- Differential treatment: direct Minor, indirect Negligible
- Civil rights: direct Negligible, indirect Minor
- Democracy: direct Negligible, indirect Negligible
- Privacy: direct Substantial, indirect Negligible
- Psychological: direct Negligible, indirect Minor
- Epistemic: direct Negligible, indirect Negligible
- Child sexual exploitation and abuse: direct Negligible, indirect Negligible
Financial
Reported: The reports mention that applicants put a lot of 'expense into the test' and hired ethics attorneys to represent them.
Directly caused: N/A
Indirectly caused: Flagged applicants incurred financial costs hiring legal counsel, such as attorneys Zavieh and Joyce, to defend against the state bar's allegations.
Inferred additional harm: Many of the 3,190 flagged applicants likely faced delayed employment and lost wages due to delayed licensing, as well as potential fees for retaking the exam.
Differential treatment
Reported: The reports explicitly mention that the cheating algorithm flagged Black and other minority applicants 'for merely existing'.
Directly caused: The AI system exhibited algorithmic bias, disproportionately flagging minority applicants based on their demographic characteristics.
Indirectly caused: N/A
Inferred additional harm: Minority applicants likely faced a higher burden of proof and systemic discrimination during the subsequent review process.
Civil rights
Reported: The reports describe concerns about the unfairness of forcing applicants to respond to vague allegations without access to the video evidence.
Directly caused: N/A
Indirectly caused: The state bar's administrative process violated basic procedural fairness and due process by denying applicants access to the evidence used against them.
Inferred additional harm: Thousands of applicants had their right to a fair administrative process compromised by the automated flagging system.
Privacy
Reported: The reports explicitly state that critics warned remote proctoring could 'invade test takers' privacy'.
Directly caused: The AI proctoring software gained deep access to applicants' personal computers, webcams, and microphones to monitor them in their homes.
Indirectly caused: N/A
Inferred additional harm: The continuous biometric and environmental surveillance of 9,301 applicants in their private homes represents a significant intrusion of privacy.
Psychological
Reported: The reports explicitly describe psychological distress, stating that flagged applicants were 'frightened' and 'angry' and experienced 'unnecessary stress'.
Directly caused: N/A
Indirectly caused: At least dozens of applicants represented by lawyers, and potentially many of the 3,190 flagged individuals, experienced severe anxiety, fear, and anger due to receiving cheating allegations.
Inferred additional harm: It is highly likely that a significant portion of the 3,190 flagged applicants suffered from acute stress and anxiety regarding their career prospects and professional reputation.
People affected
- Occurrences reported: 1
- People reportedly harmed: 3190
- People reportedly exposed: 9301
Potential causes
Management
- Over-reliance on Automated Flags: Management assumed the AI could accurately detect academic dishonesty.
- Inadequate Resource Planning: Bar was unprepared to review the massive volume of flagged videos.
- Lack of Accountability: Bar shifted the burden of proof to applicants rather than verifying.
Technology
- Biased Flagging Algorithm: Algorithm flagged Black and diverse applicants at disproportionate rates.
- Overly Sensitive AI Proctoring: System flagged benign behaviors like gazing off-screen as cheating.
- Technical Software Issues: 15 percent of test takers experienced technical issues on exam day.
Data Inputs
- Biased Training Data: Lack of diverse representation led to false cheating flags for minorities.
- Video and Audio Infractions: Environmental inputs like silence or eye gaze triggered automatic flags.
- Prohibited Items in View: Presence of food or digital clocks in video feed triggered alerts.
Human Factors
- Natural Test-Taking Behaviors: Applicants naturally looked away or remained silent, triggering flags.
- Lack of Proctor Guidance: No live proctors to prevent minor rule infractions before the exam.
- Remote Environment Confusion: Applicants took the test in home environments with unavoidable distractions.
Process and Methods
- Lack of Human Pre-Verification: Examiners sent notices without reviewing videos to verify AI flags.
- Vague Chapter 6 Notices: Notices contained generic descriptions of alleged violations.
- Denial of Evidence Access: Applicants had to respond to allegations without seeing their videos.
Regulatory Environment
- Pandemic Emergency Mandate: COVID-19 forced a rapid, unregulated shift from in-person to online exams.
Information quality
- Classification confidence: High
- Reason for confidence: The reports provide precise statistics (9,301 test-takers, 3,190 flagged) and clear descriptions of the AI system's role, the developer (ExamSoft), and the specific harms (stress, legal costs, potential licensing delays) experienced by the affected parties.
- Ambiguities identified: The exact algorithmic criteria used by ExamSoft to flag applicants are not fully detailed, and the extent of the racial bias mentioned in Report 1 is not quantified.
- Alternative interpretations: None. The role of the AI in generating the flags that led to the administrative actions is clear and undisputed.
An AI-powered proctoring system unintentionally flagged over a third of California bar exam takers for cheating due to high sensitivity and potential demographic bias. While causing significant stress, financial costs, and privacy concerns for thousands of applicants, the incident is a domestic administrative failure with negligible direct impact on national security.
- Overall national security impact: Minor
- Response level: Moderate
- Scope: Single nation
- Primary target: United States
- Alleged perpetrator: Unknown
Threat characteristics
- Imminence: Long-term. The event was a temporary administrative issue that does not pose an ongoing or immediate national security crisis.
- Autonomy: Human-supervised. The AI software autonomously flagged suspicious behavior, but humans (the state bar examiners) were responsible for reviewing the flags and making final licensing decisions.
- Novelty: Evolved capability. Represents an evolved use case of computer vision and biometric monitoring scaled rapidly for high-stakes professional licensing during a pandemic.
Impact by dimension
- Physical security: Negligible. The incident involved online exam proctoring and did not pose any threat to physical systems, critical infrastructure, or human safety.
- Information security: Negligible. No compromise of intelligence capabilities, classified models, or foreign disinformation campaigns occurred during this academic testing incident.
- Sovereignty: Negligible. Although the incident disrupted state-level licensing for lawyers, it did not threaten national sovereignty, state authority, or core constitutional processes.
- Economic security: Negligible. While individual test-takers faced legal fees and potential career delays, there was no threat to national financial stability or strategic technological competitive advantage.
- Societal stability: Minor. The system caused widespread psychological distress, potential demographic bias, and privacy concerns for over 3,000 citizens, representing a minor domestic human rights concern.