ChatGPT Abused to Develop Malicious Softwares

Following the release of ChatGPT and Codex, cybersecurity researchers and threat actors demonstrated that these models could be used to generate functional malicious code, phishing emails, and malware components. The tools were found to lower the barrier to entry for cybercrime, allowing individuals with little to no coding experience to develop and iterate on malicious software, despite OpenAI's safety guardrails.

OpenAI's ChatGPT was reportedly abused by cyber criminals including ones with no or low levels of coding or development skills to develop malware, ransomware, and other malicious softwares.

Source: AI Incident Database

Risk classification

  • Primary risk domain: 4 Malicious actors
  • Primary risk subdomain: 4.2 Cyberattacks, weapon development or use, and mass harm

The incident involves cybercriminals using ChatGPT to develop cyber weapons, specifically coding malware such as infostealers, backdoors, and ransomware concepts.

Additional risk subdomains

  • 4.3 Fraud, scams, and targeted manipulation: Threat actors and researchers used the AI to generate highly convincing phishing emails and plan automated romance scams using fake personas.

Causal factors

  • Entity: Human
  • Intent: Intentional
  • Timing: Post-deployment

The risk is driven by human cybercriminals intentionally exploiting and manipulating the deployed AI system to generate malicious software and phishing materials.

EU AI Act risk tier

  • Risk tier: 3 Limited Risk

Limited Risk: The reports describe ChatGPT as an 'AI chatbot' and 'natural language processing tool' generating text and code, which falls under the transparency obligations for chatbots and generative AI.

AI system and alleged parties

  • AI system: ChatGPT, Codex (OpenAI)
  • AI purpose: Code Generation; Writing Assistant
  • Behaviour type: Assistant
  • Alleged developer: OpenAI
  • Alleged deployer: OpenAI
  • Alleged harmed parties: internet users

Harm severity

Highest direct severity in any category: Substantial. Severity is scored from Negligible to Catastrophic in each harm category, for harm the reports describe as caused directly or indirectly by the AI system.

  • Physical: direct Negligible, indirect Negligible
  • Infrastructure: direct Negligible, indirect Negligible
  • Property: direct Negligible, indirect Negligible
  • Financial: direct Negligible, indirect Negligible
  • Environmental: direct Negligible, indirect Negligible
  • Malicious content: direct Minor, indirect Minor
  • Differential treatment: direct Negligible, indirect Negligible
  • Civil rights: direct Negligible, indirect Negligible
  • Democracy: direct Negligible, indirect Negligible
  • Privacy: direct Negligible, indirect Negligible
  • Psychological: direct Negligible, indirect Negligible
  • Epistemic: direct Negligible, indirect Negligible
  • Child sexual exploitation and abuse: direct Negligible, indirect Negligible

Malicious content

Reported: The report explicitly describes the generation of malicious code, phishing emails, and scam scripts directly caused by the incident.

Directly caused: ChatGPT directly generated functional malicious Python, Java, and VBA code, as well as highly convincing phishing emails when prompted by researchers and forum users.

Indirectly caused: The dissemination of these AI-generated code templates on underground forums facilitated the spread of malware development techniques to low-skilled actors.

Inferred additional harm: It is highly likely that many more malicious scripts and phishing emails have been generated and distributed than those specifically documented by Check Point Research.

People affected

  • Occurrences reported: 1

Potential causes

Management

  • Rapid public release of tool: Broad public release occurred before establishing foolproof safety controls.
  • No penalties for policy violations: No immediate consequences are enforced for violating terms of service.

Technology

  • Dual-use nature of computer code: Benign programming functions can easily be modified for malicious purposes.
  • Imperfect safety guardrails: Filters can be bypassed by rephrasing prompts or using roleplay tactics.
  • Low barrier to entry: Conversational interface enables non-technical users to generate malware.

Data Inputs

  • Training on public code repositories: Vast training data includes public code containing vulnerabilities and exploits.
  • Abuse of malware research text: Users input published research papers to recreate specific malware strains.

Human Factors

  • Malicious intent of threat actors: Cybercriminals actively abuse the chatbot to develop offensive hacking tools.
  • Low technical skill of users: Allows script kiddies with zero coding skills to launch cyberattacks.
  • Human vulnerability to deception: Targets are easily fooled by highly convincing, error-free phishing emails.

Process and Methods

  • Iterative prompting techniques: Attackers refine basic outputs step-by-step to create complex malware.
  • Lack of user intent verification: No process exists to verify whether code requests are for benign use.

Regulatory Environment

  • Absence of AI safety regulations: No legal framework forces developers to secure models against cybercrime.
  • Lack of regulatory enforcement: No oversight body exists to penalize developers for abuse of AI tools.

Information quality

  • Classification confidence: High
  • Reason for confidence: The reports are highly consistent, drawing from a detailed primary source (Check Point Research) with concrete examples of generated code (Python infostealers, Java downloaders, VBA macros). The technical capabilities and limitations are well-documented across multiple reputable tech journalism outlets.
  • Ambiguities identified: Whether any of the generated malware has been successfully deployed in active, real-world attacks remains unconfirmed.

Cybersecurity researchers and threat actors demonstrated that ChatGPT and Codex could be leveraged to generate functional malicious code and phishing emails. While lowering the barrier to entry for low-skilled cybercriminals, the incident represents an evolved capability for cyber threat generation rather than an active national security crisis.

  • Overall national security impact: Minor
  • Response level: Moderate
  • Scope: Multiple nations
  • Primary target: No clear primary
  • Other affected: Multiple nations
  • Alleged perpetrator: Unknown

Threat characteristics

  • Imminence: Long-term. Represents an ongoing strategic concern regarding the democratization of cyberweapon development rather than an active, immediate crisis.
  • Autonomy: Human-controlled. The AI acts as a writing and coding assistant; humans must actively prompt the system, bypass guardrails, and deploy the generated code.
  • Novelty: Evolved capability. Significant advancement in how existing cyber threats like malware and phishing are authored, lowering the technical barrier to entry.

Impact by dimension

  • Physical security: Negligible. No physical systems, critical infrastructure, or human safety were impacted or targeted in the reported incident.
  • Information security: Minor. Demonstrates potential for AI-enhanced phishing and malware generation, but no active state-sponsored intelligence compromise or mass disinformation campaigns were reported.
  • Sovereignty: Negligible. No disruption to state authority, electoral systems, or core government operations was reported.
  • Economic security: Minor. Lowered the barrier to entry for cybercrime, enabling novice actors to develop malware, though no active widespread financial damage was confirmed in the report.
  • Societal stability: Negligible. No mass surveillance, systematic discrimination, or large-scale threats to social cohesion or civil liberties were reported.
Explore in the interactive Incident Tracker