Flawed AI in Google Search Reportedly Misinforms about Geography

Google's search engine and its experimental 'Search Generative Experience' (SGE) were found to be providing factually incorrect information to users, such as claiming no African countries begin with the letter 'K'. The errors stemmed from the systems scraping and paraphrasing flawed AI-generated content from the web, as well as inherent failures in the models' ability to process logic and alphabetization tasks. This incident highlights the risks of integrating generative AI into search platforms, where hallucinations are presented as authoritative facts.

Google's search AI erroneously claimed no African country begins with 'K', along with various other geography-and-letter-based questions, misguiding users with a flawed featured snippet. Originating from ChatGPT-written posts and inaccurately scraped by Google, this incident highlights issues in AI-generated content and misinformation in search results, compromising Google's reliability as an information source.

Source: AI Incident Database

Risk classification

  • Primary risk domain: 3 Misinformation
  • Primary risk subdomain: 3.1 False or misleading information

The AI systems generated and spread incorrect geographical and biographical information, presenting hallucinations as authoritative facts to users.

Additional risk subdomains

  • 7.3 Lack of capability or robustness: The models exhibited a fundamental lack of capability in handling basic logical, alphabetical, and spatial reasoning tasks.

Causal factors

  • Entity: AI
  • Intent: Unintentional
  • Timing: Post-deployment

The factual errors and hallucinations were caused by the AI systems' processing failures after being deployed to the public.

EU AI Act risk tier

  • Risk tier: 3 Limited Risk

Limited Risk: The systems involved are chatbots and generative AI search features. According to the EU AI Act, these systems pose a moderate risk and are subject to transparency obligations, such as ensuring users are aware they are interacting with an AI.

AI system and alleged parties

  • AI system: ChatGPT, SGE (Google, OpenAI)
  • AI purpose: Content Search; Question Answering
  • Behaviour type: Assistant
  • Alleged developer: Google, ChatGPT
  • Alleged deployer: Google
  • Alleged harmed parties: General public

Harm severity

Highest direct severity in any category: Severe. Severity is scored from Negligible to Catastrophic in each harm category, for harm the reports describe as caused directly or indirectly by the AI system.

  • Physical: direct Negligible, indirect Negligible
  • Infrastructure: direct Negligible, indirect Negligible
  • Property: direct Negligible, indirect Negligible
  • Financial: direct Negligible, indirect Negligible
  • Environmental: direct Negligible, indirect Negligible
  • Malicious content: direct Negligible, indirect Negligible
  • Differential treatment: direct Negligible, indirect Negligible
  • Civil rights: direct Negligible, indirect Negligible
  • Democracy: direct Negligible, indirect Negligible
  • Privacy: direct Negligible, indirect Negligible
  • Psychological: direct Negligible, indirect Negligible
  • Epistemic: direct Minor, indirect Substantial
  • Child sexual exploitation and abuse: direct Negligible, indirect Negligible

Epistemic

Reported: The report explicitly describes epistemic harm, where Google Search and SGE generated and spread false and misleading information.

Directly caused: The AI systems directly generated and displayed false claims, such as asserting there are no African countries starting with 'K' and providing incorrect alphabetical lists of European countries.

Indirectly caused: The search engine elevated ChatGPT-generated nonsense from third-party blogs to featured snippets, polluting the information ecosystem.

Inferred additional harm: It is highly likely that many users who saw these featured snippets believed the false information without clicking through to verify, leading to widespread incorrect beliefs.

People affected

  • Occurrences reported: 1
  • People reportedly exposed: 10000

Potential causes

Management

  • Layoffs in Fact-Checking Teams: Layoffs in Google News division reduced resources for manual quality efforts.
  • Prioritizing Speed Over Quality: Deploying experimental SGE features to the public despite known inaccuracies.

Technology

  • LLM Alphabetization Limits: LLMs struggle to parse spelling, alphabetization, and letter-based queries.
  • SGE Hallucination Tendency: Generative AI in search synthesizes incorrect facts and presents them as truth.
  • Flawed Snippet Algorithms: Algorithms pull and elevate unverified, low-quality text from the open web.

Data Inputs

  • Ingestion of AI-Generated Spam: Crawlers scraped a blog post quoting ChatGPT nonsense, treating it as fact.
  • Unverified Source Material: System relied on a Hacker News thread and a personal obituary for facts.
  • Circular Web Data Pollution: AI-generated content is indexed and then re-digested by search LLMs.

Human Factors

  • Superficial User Verification: Users may read SGE results superficially instead of clicking sources to verify.

Process and Methods

  • Insufficient Fact-Checking: Lack of automated or manual fact-checking to verify SGE and snippet outputs.
  • No Manual Correction Policy: Google does not manually fix untrue snippets unless they cause direct harm.
  • Inadequate RAG Verification: Retrieval-augmented generation failed to cross-check and weed out false claims.

Information quality

  • Classification confidence: High
  • Reason for confidence: The reports provide clear, first-hand accounts and screenshots of the AI systems' outputs, along with statements from Google representatives. The nature of the failure (hallucination and scraping of AI-generated junk) is well-documented and unambiguous.
  • Ambiguities identified: None significant; the factual errors made by the AI are clear and verifiable.
  • Alternative interpretations: The incident could be viewed purely as a search engine indexing failure, but the generative AI aspect (SGE and ChatGPT) makes it a clear AI safety and robustness issue.

Google Search and its experimental Search Generative Experience (SGE) propagated factual errors by scraping and elevating ChatGPT-generated hallucinations. While demonstrating vulnerabilities in AI-driven information retrieval and degrading the online information ecosystem, the national security implications remain minor and confined to general epistemic trust.

  • Overall national security impact: Minor
  • Response level: Moderate
  • Scope: Multiple nations
  • Primary target: No clear primary
  • Other affected: Global users
  • Alleged perpetrator: Unknown

Threat characteristics

  • Imminence: Long-term. Represents an ongoing systemic challenge with AI hallucinations and information quality rather than an immediate national security crisis.
  • Autonomy: Human-supervised. The AI systems autonomously scraped, summarized, and displayed information, though Google maintains overarching algorithmic and manual suppression controls.
  • Novelty: Evolved capability. While LLM hallucinations are well-established, the integration of generative AI into global search engines represents an evolved delivery method.

Impact by dimension

  • Physical security: Negligible. No physical systems, infrastructure, or human safety were threatened or impacted by these search engine hallucinations.
  • Information security: Minor. While not an active hostile campaign, the systemic elevation of false AI-generated content degrades the broader information ecosystem and trust in authoritative sources.
  • Sovereignty: Negligible. No state authority, border control, or core government decision-making processes were compromised by the search errors.
  • Economic security: Negligible. The incident represents a product quality and reputational issue for Google rather than a strategic threat to national economic stability or critical technology theft.
  • Societal stability: Negligible. The epistemic harm was limited to minor factual errors and did not result in mass surveillance, systematic discrimination, or civil unrest.
Explore in the interactive Incident Tracker