Google Gemini CLI Reportedly Deletes User Files After Misinterpreting Command Sequence

AI coding assistants from Google and Replit caused significant data loss by misinterpreting system commands and failing to verify operations. In the Gemini incident, the AI hallucinated a successful directory creation, leading to the overwriting and deletion of files. These events highlight critical safety risks in 'vibe coding' tools that lack robust error-checking.

Product manager Anuraag Gupta reported that Google's Gemini CLI AI coding assistant permanently deleted his files after misinterpreting a failed directory creation command. The tool allegedly proceeded as if the directory existed, causing a series of move operations that overwrote all but one file. Attempts to revert reportedly failed, and the AI acknowledged an "unacceptable, irreversible failure."

Source: AI Incident Database

Risk classification

  • Primary risk domain: 7 AI system safety, failures, & limitations
  • Primary risk subdomain: 7.3 Lack of capability or robustness

The incident represents a clear lack of capability and robustness in the AI systems, which hallucinated command success and failed to perform basic read-after-write verification, leading to catastrophic data loss.

Additional risk subdomains

  • 5.1 Overreliance and unsafe use: Users engaged in 'vibe coding' placed high trust in AI assistants to execute terminal commands directly on their file systems and databases without manual verification.

Causal factors

  • Entity: AI
  • Intent: Unintentional
  • Timing: Post-deployment

The data loss was caused by the Gemini CLI and Replit AI systems post-deployment, which unintentionally executed destructive file operations due to hallucinations and a lack of verification.

EU AI Act risk tier

  • Risk tier: 4 Minimal or No Risk

Minimal or No Risk: The AI systems are general-purpose developer tools and coding assistants, which do not fall under the prohibited or high-risk categories defined in the Act.

AI system and alleged parties

  • AI system: Gemini CLI (Google)
  • AI purpose: Code Generation; Writing Assistant
  • Behaviour type: Agent
  • Alleged developer: Google
  • Alleged deployer: Google
  • Alleged harmed parties: Users of Gemini CLI, Anuraag Gupta

Harm severity

Highest direct severity in any category: Negligible. Severity is scored from Negligible to Catastrophic in each harm category, for harm the reports describe as caused directly or indirectly by the AI system.

  • Physical: direct Negligible, indirect Negligible
  • Infrastructure: direct Negligible, indirect Negligible
  • Property: direct Negligible, indirect Negligible
  • Financial: direct Negligible, indirect Negligible
  • Environmental: direct Negligible, indirect Negligible
  • Malicious content: direct Negligible, indirect Negligible
  • Differential treatment: direct Negligible, indirect Negligible
  • Civil rights: direct Negligible, indirect Negligible
  • Democracy: direct Negligible, indirect Negligible
  • Privacy: direct Negligible, indirect Negligible
  • Psychological: direct Negligible, indirect Negligible
  • Epistemic: direct Negligible, indirect Negligible
  • Child sexual exploitation and abuse: direct Negligible, indirect Negligible

People affected

  • Occurrences reported: 2
  • People reportedly harmed: 2
  • People reportedly exposed: 2

Potential causes

Management

  • Prioritizing Speed Over Safety: Releasing CLI tools to compete with rivals without robust safety checks.
  • Inadequate Risk Assessment: Failure to assess dangers of giving AI agents direct terminal write access.

Technology

  • Lack of Verification Checks: The AI failed to confirm if the mkdir command succeeded before moving files.
  • Hallucination of File Operations: The AI acted on the false premise that a directory was successfully created.
  • Implicit Self-Trust: The AI trusted its own actions without performing read-after-write validation.

Human Factors

  • Overreliance on AI Automation: The user trusted the AI agent to manage files without manual oversight.
  • Vibe Coding Trend Adoption: Users embrace fluid, vibe-based coding without verifying critical steps.

Process and Methods

  • Lack of Read-After-Write Check: No process existed to verify directory creation before executing file moves.
  • Inadequate System Safeguards: The agent lacked guardrails to prevent destructive OS-level commands.

Information quality

  • Classification confidence: High
  • Reason for confidence: The report provides clear, first-hand post-mortem details of the Gemini CLI failure and references a verified public incident with the Replit agent. The technical steps of the failure (mkdir failure, Windows move command behavior, sequential overwriting) are explicitly documented.
  • Ambiguities identified: None significant; the technical mechanism of the file deletion is clearly explained.
  • Alternative interpretations: The Replit incident could be interpreted as a user-error in database permissions, but the report attributes it to AI hallucination and lack of guardrails.

AI coding assistants from Google and Replit caused localized data loss and temporary database deletion due to hallucinations and a lack of verification. While highlighting critical safety risks in developer tools and potential software supply chain vulnerabilities, the incident has negligible direct national security implications.

  • Overall national security impact: Minor
  • Response level: Moderate
  • Scope: Single nation
  • Primary target: United States
  • Alleged perpetrator: Unknown

Threat characteristics

  • Imminence: Long-term. Represents an ongoing strategic concern regarding AI code generation safety and developer overreliance, rather than an active national security crisis.
  • Autonomy: Human-supervised. The AI agent executed multi-step file and database operations autonomously based on high-level human instructions, but operated within a user-initiated session.
  • Novelty: Evolved capability. Represents a significant evolution in AI developer tools having direct system write access, moving from passive code generation to active, unverified system execution.

Impact by dimension

  • Physical security: Negligible. The incident involved localized digital data loss on developer machines and a commercial database, with no impact on physical systems, critical infrastructure, or human safety.
  • Information security: Negligible. No classified systems, intelligence capabilities, or systematic information warfare operations were involved or compromised.
  • Sovereignty: Negligible. The incident affected private individuals and commercial entities, with no disruption to state authority, electoral systems, or core government operations.
  • Economic security: Minor. Represented localized disruption and temporary data loss for commercial developers, highlighting software supply chain vulnerabilities, but had negligible impact on national economic or strategic technological security.
  • Societal stability: Negligible. No implications for civil liberties, mass surveillance, or societal stability; the impact was confined to individual developers and a business.
Explore in the interactive Incident Tracker