Perplexity AI Grounding Analysis: Manufactured Sources and AI-Generated Buying Guides

AI Retrieval Systems are Grounding Recommendations in Low-Authority Domains

Research into web-grounded AI models reveals a significant reliance on low-authority and machine-generated content for product recommendations. In a test of 380 software categories using perplexity/sonar and perplexity/sonar-pro, 83.2% of the 7,534 citations pointed to domains ranked worse than #100,000 in the Tranco top-1M list or domains that were not ranked in the top million at all.

Distribution of Citation Authority

The retrieval layer for Perplexity's models shows a marked lack of concentration among high-authority sites. While the top ten most-cited domains account for 17.3% of citations, the remaining four-fifths of the evidence base is composed of long-tail sources. Specifically, 36.5% of the cited domains do not appear in the Tranco top million. These unranked domains are significantly newer, with a median first Wayback Machine capture in 2020, compared to 2011 for ranked domains.

Domain Citations Share Tranco rank
g2.com 291 3.86% 4,027
reddit.com 261 3.46% 105
guideflow.com 194 2.57% 177,039
gartner.com 158 2.10% 1,766
zapier.com 82 1.09% 2,919

The Rise of "Grounding Pages" for AI SEO

A network of three domains—wifitalents.com, worldmetrics.org, and gitnux.org—was identified as a primary source of manufactured content designed specifically for AI retrieval systems. These sites, all registered between December 2023 and May 2024, share identical Cloudflare nameservers, page templates, and navigation structures.

Scale of Machine-Generated Content

Between them, these three sites published 215,128 machine-generated best <category> pages. The scale of this operation far exceeds the number of existing software categories, indicating a programmatic approach to capturing AI search traffic.

Targeting the Retrieval Layer

Unlike traditional SEO, which targets human readers, these sites explicitly target the AI "grounding" process. worldmetrics.org and gitnux.org utilize HTML titles such as "Facts & Grounding Page" and meta descriptions that describe themselves as "machine-readable record[s] of verified facts." This suggests a shift toward "AI Engine Optimization" (AEO), where content is structured specifically to be ingested by LLM retrieval layers.

Inconsistencies in AI-Generated Recommendations

Despite presenting as authoritative research firms, the network of manufactured sites provides contradictory rankings for the same software categories. For the category "project estimation software," the three sites provided three different top-ranked tools, despite using identical templates and claiming "AI-verified" or "Expert reviewed" processes.

Furthermore, the retrieval process can lead to hallucinated or dangerous links. In the study's sample of 1,502 vendor homepages supplied by the models:

  • 1.1% were gone or unreachable.
  • 6.1% redirected to different domains.
  • Some citations led to unrelated sites, such as an Indonesian online-gambling portal and a Monaco hotel and casino group, when the model was asked for research data platforms and data quality tools.

Community Insights and Counterpoints

Discussion among technical users suggests that this phenomenon is part of a broader trend of "AI slop" and circular reasoning in the AI ecosystem.

"The search engine is now the citation, and the citation is a page that exists to be cited. Nobody in that loop has read anything, and it still works."

Key concerns raised by the community include:

  • Model Preference for AI Content: Some users report that LLMs tend to favor AI-generated passages over human-written ones, creating a feedback loop where models train on and cite their own regurgitations.
  • Source Skepticism: There is a perceived lack of "source skepticism" in current agentic workflows, where models fail to consider the motive of the published information, making them vulnerable to "pay-for-play" verification systems.
  • Authenticity of the Report: Some HN commenters questioned the authenticity of the research itself, suggesting the report may be part of the same AI-generated content strategy it describes, noting that the reporting site (trellner.com) appeared to follow similar AI-generated patterns.

Methodology and Scope

The study focused exclusively on Perplexity AI's sonar and sonar-pro models via OpenRouter. It utilized 380 buyer-intent categories and 760 total calls. The researchers noted that the results reflect a snapshot of a retrieval index and do not necessarily represent the performance of other AI search engines like ChatGPT, Gemini, or Copilot.

Sources

Related