Yale Student Suspended and Given F After GPTZero Allegedly Misclassified Exam as AI-Generated

A Yale Executive MBA student was suspended for a year and given an F after an AI detection tool, GPTZero, flagged his exam as AI-generated, leading to a major lawsuit alleging bias against non-native English speakers.

In 2024, Yale University reportedly used GPTZero as one factor in investigating whether Executive MBA student Thierry Rignol used AI on an exam. Rignol reportedly denied cheating and alleged the detector falsely classified his writing as AI-generated. Yale reportedly later suspended him for a year and gave him an F, while citing additional evidence and his conduct during the investigation. Rignol has challenged the process in federal court.

Source: AI Incident Database

Risk classification

  • Primary risk domain: 1 Discrimination & Toxicity
  • Primary risk subdomain: 1.3 Unequal performance across groups

The incident centers on the AI detector's alleged bias against non-native English speakers, where the tool's unequal accuracy across demographic groups led to a false cheating accusation and severe academic penalties.

Additional risk subdomains

  • 7.3 Lack of capability or robustness: The AI detection tool GPTZero demonstrated a severe lack of robustness and reliability by generating false positives, including flagging 30-year-old human-written texts as 100% AI-generated.

Causal factors

  • Entity: AI
  • Intent: Unintentional
  • Timing: Post-deployment

The incident was triggered by the biased and unreliable output of the deployed GPTZero AI detection tool, which unintentionally flagged a non-native English speaker's writing as AI-generated.

EU AI Act risk tier

  • Risk tier: 2 High Risk

Risk Level 2: High Risk. The AI system was used in an educational setting to assess student work and determine disciplinary action, which directly affects access to education as defined under High Risk systems ('Educational and vocational training systems affecting access to education').

AI system and alleged parties

  • AI system: ChatGPT, GPTZero (GPTZero, OpenAI)
  • AI purpose: Cheating Detection; Content Verification
  • Behaviour type: Tool
  • Alleged developer: GPTZero, AI text detection system developers
  • Alleged deployer: Yale University, Universities, Professors, Institutions of higher education, Educational communities
  • Alleged harmed parties: University students, Thierry Rignol, Students, Non-native English speakers

Harm severity

Highest direct severity in any category: Minor. Severity is scored from Negligible to Catastrophic in each harm category, for harm the reports describe as caused directly or indirectly by the AI system.

  • Physical: direct Negligible, indirect Negligible
  • Infrastructure: direct Negligible, indirect Negligible
  • Property: direct Negligible, indirect Negligible
  • Financial: direct Negligible, indirect Minor
  • Environmental: direct Negligible, indirect Negligible
  • Malicious content: direct Negligible, indirect Negligible
  • Differential treatment: direct Negligible, indirect Minor
  • Civil rights: direct Negligible, indirect Negligible
  • Democracy: direct Negligible, indirect Negligible
  • Privacy: direct Negligible, indirect Negligible
  • Psychological: direct Negligible, indirect Minor
  • Epistemic: direct Negligible, indirect Minor
  • Child sexual exploitation and abuse: direct Negligible, indirect Negligible

Financial

Reported: The report explicitly describes financial losses, noting that Rignol paid $208,500 in tuition and sued for past and future economic losses.

Directly caused: N/A

Indirectly caused: Average financial loss per occurrence: $208,500. Rignol suffered indirect financial harm from the temporary loss of his tuition value due to a one-year suspension, delayed career opportunities, and substantial legal fees from his ongoing federal lawsuit.

Inferred additional harm: Legal fees for a federal lawsuit spanning over a year with 125 docket entries are estimated to exceed $100,000.

Differential treatment

Reported: The report explicitly describes differential treatment of a non-native English speaker caused directly by the AI system's bias.

Directly caused: The GPTZero tool exhibited differential accuracy by flagging the formal, structured writing style of a non-native English speaker as AI-generated text.

Indirectly caused: Yale's disciplinary process relied on the biased AI tool's output to target and penalize Rignol, resulting in unequal academic outcomes compared to native English-speaking peers.

Inferred additional harm: N/A

Psychological

Reported: The report explicitly describes psychological harm, noting that Rignol sued Yale for emotional distress.

Directly caused: N/A

Indirectly caused: Rignol experienced significant emotional distress, anxiety, and reputational damage due to being falsely accused of cheating and suspended from his program.

Inferred additional harm: N/A

Epistemic

Reported: The report explicitly describes epistemic harm in the form of false accusations and fabricated cheating claims generated by the AI tool.

Directly caused: GPTZero generated false positive results, incorrectly identifying Rignol's exam answers and historical texts by Yale administrators as AI-generated.

Indirectly caused: The false positive outputs misled Yale administrators, corrupting the academic evaluation process and leading to an unjustified suspension.

Inferred additional harm: N/A

People affected

  • Occurrences reported: 1
  • People reportedly harmed: 1
  • People reportedly exposed: 1

Potential causes

Management

  • Poor Academic Risk Assessment: Management failed to assess risks of using biased AI detection tools.
  • Escalation of Disciplinary Pressure: Administrators used high-pressure tactics to force a confession.

Technology

  • Unreliable AI Detection Tools: GPTZero and similar tools are highly unreliable for detecting AI text.
  • Detector Bias Against Non-Native Writers: Detectors mistake formal non-native English prose for AI-generated text.
  • High False Positive Rates: Scans of human-written historical texts falsely flagged as 100% AI.

Data Inputs

  • Linguistic Patterns of Test Text: Formal, structured, and long test prose triggered detector flags.
  • Overlap with ChatGPT Output: Student answers had substantial overlap with ChatGPT generated outputs.

Human Factors

  • Over-reliance on AI Detectors: Teaching staff relied heavily on automated tools to flag cheating.
  • Communication Failures: Student withheld Pages file because Yale asked for a Word file.
  • Misinterpretation of Writing Style: Professors mistook high-quality, polished writing for AI output.

Process and Methods

  • Infeasible AI Policing Methods: Using automated detectors to police academic integrity is highly flawed.
  • Lack of Clear AI Verification Steps: No robust process to verify AI use beyond detector tools.
  • Inadequate Due Process: Student suspended for lack of candor without a formal charge.

Regulatory Environment

  • Lack of Standardized AI Rules: No clear institutional standards on validating AI detection tool use.

Information quality

  • Classification confidence: High
  • Reason for confidence: The report provides a highly detailed account of the dispute between Thierry Rignol and Yale University, including specific dates, the name of the AI tool (GPTZero), the academic and financial consequences, and the arguments from both sides. The role of the AI detector as a trigger for the incident is clear, and the limitations of the tool are well-documented.
  • Ambiguities identified: It remains legally unresolved whether Rignol actually used AI or if the tool's output was entirely a false positive, though the university's final punishment was officially for 'not being forthcoming' rather than solely AI use.
  • Alternative interpretations: The incident could be interpreted primarily as an administrative and disciplinary dispute regarding student cooperation rather than an AI safety failure, given that the suspension was triggered by Rignol's failure to provide his Pages file.

A Yale EMBA student was suspended following a false positive flag from the AI detector GPTZero. The incident highlights ongoing issues with AI detection reliability and demographic bias, but it is a localized academic and legal dispute with negligible national security implications.

  • Overall national security impact: Negligible
  • Response level: Minor
  • Scope: Single nation
  • Primary target: United States
  • Alleged perpetrator: Unknown

Threat characteristics

  • Imminence: Long-term. This incident represents a localized, resolved academic disciplinary action with no imminent threat to national security.
  • Autonomy: Human-controlled. The AI detection tool was used to assist decision-making, with Yale administrators and the Honor Committee retaining full control over the final disciplinary actions.
  • Novelty: Established threat. False positives and demographic biases in AI text detection tools are well-documented issues that have occurred in numerous academic settings.

Impact by dimension

  • Physical security: Negligible. The incident involves an academic integrity dispute at a university and has no impact on physical systems, critical infrastructure, or human safety.
  • Information security: Negligible. There is no compromise of intelligence capabilities, classified models, or state-sponsored information warfare indicated in this incident.
  • Sovereignty: Negligible. This is a localized academic disciplinary matter with no threat to state authority, electoral systems, or core government operations.
  • Economic security: Negligible. The financial impact is limited to individual tuition loss and legal fees, presenting no threat to national economic stability or strategic industries.
  • Societal stability: Negligible. While the incident involves allegations of bias against non-native English speakers, it remains a localized academic dispute rather than a large-scale threat to societal stability.
Explore in the interactive Incident Tracker