
Summary
- Trellner queried Perplexity's sonar and sonar-pro models across 380 software categories, analyzing 7,534 citations
- 59.8% of citations pointed to domains outside the top 100,000 by traffic, and three sites had auto-generated 215,128 pages
- worldmetrics.org and gitnux.org titled their homepages "Facts & Grounding Page," directly targeting the AI grounding step
When you ask an AI which software to buy, what does it actually rely on for its answer? According to a study released September 2, 2026 by research outlet Trellner, many of the documents Perplexity cited when answering questions like "what's the best CRM software" turned out to be sites with almost no visitors. Of the 7,534 citations gathered across 380 software categories, 59.8% pointed to domains ranked outside the top 100,000 by web traffic, and three of those sites had auto-generated more than 215,000 pages, branding themselves internally as "Facts & Grounding Pages."
380 Categories, 760 Queries
On September 2, 2026, Trellner's researchers accessed Perplexity's sonar and sonar-pro models through OpenRouter and asked each of them about 380 purchase-intent categories, ranging from "CRM software" to "museum collection management software." That came to 760 total calls — one per category per model — and every response returned five recommended products along with the URLs the model had actually pulled from its search. The researchers say they locked in the category list before seeing any results and didn't revise it afterward.
That produced 7,534 citations across 2,055 distinct domains. The researchers cross-checked those domains against the Tranco list, a web traffic ranking, and against the Internet Archive's Wayback Machine. Among ranked domains, the median citation ranked 71,611th by traffic; 59.8% of all citations fell outside the top 100,000, and 23.4% didn't even crack the top million. For comparison, Wikipedia was cited just 3 times out of the 7,534.
Google wasn't part of this study. The researchers explained that routing Gemini-family models through OpenRouter for grounding runs the query through OpenRouter's own search plugin rather than Google's native search.
A Company Not Even Competing Became the No. 3 Source
The most striking case is guideflow.com. It's a company that sells interactive product demos — not a review site, not a directory — and it doesn't actually compete in any of the 380 categories studied. Even so, its blog was cited 194 times across 96 categories, a quarter of the total, making it the third most-cited source overall — ahead of market research firm Gartner. Each of the 96 categories pulled a different blog post, and the company's sitemap lists 3,351 blog URLs, 2,176 of them distinct posts. The same blog turned up as supporting evidence across wildly unrelated fields — "3D rendering software," "IVR software," "RFID software," "architecture practice management software." The researchers note that running a large content-marketing blog isn't unusual in itself, but they flagged the fact that posts from a company with no stake in these markets ended up backing purchase recommendations.
"Facts & Grounding Page" — The Story Behind 215,000 Pages
The more striking case is a trio: worldmetrics.org, gitnux.org, and wifitalents.com. They were cited 60, 50, and 71 times respectively — 181 citations total across 41 categories, only 2.4% of all citations, which isn't huge on its own. But the researchers say the evidence points to all three being run by the same operator. All three domains were registered through Namecheap between December 2023 and May 2024, all delegate their DNS to the same two Cloudflare nameservers (pam.ns.cloudflare.com and sean.ns.cloudflare.com), and even their navigation menus follow an identical template. Each site carries exactly six blog posts, and every single one of them talks about the other two brands plus a fourth brand, zipdo.co. zipdo.co uses the same nameservers too, and its own homepage carries the identical title: "Facts & Grounding Page."
The three sitemaps list 103,578, 107,083, and 105,541 URLs respectively, and within those, auto-generated buying-guide pages in the format "/best/(category)-software/" number 70,731, 71,684, and 72,713. That adds up to 215,128 pages — and there simply aren't that many software categories in the world.
The HTML titles on worldmetrics.org's and gitnux.org's homepages literally read "(brand name) — Facts & Grounding Page," and their meta descriptions are nearly identical apart from the brand name.
To unpack that: grounding is the step where an AI searches actual web documents for supporting material before it writes an answer. This study tracked exactly what Perplexity pulls during that step, and it turned up sites that clearly built their page titles and descriptions for the AI passing through that step — not for human visitors.
On that same page, worldmetrics.org advertises custom market research "from €5,000," finished reports "from €499," and vendor-selection services "from €2,500." In other words, the paid services sit right above the auto-generated "best of" lists that the AI is citing.
Same Question, Different Answers — Checking With Project Estimation Software
The researchers pulled the "project estimation software" category page from all three sites and compared them directly. Each page had its rankings encoded in JSON-LD, so there was no need for interpretation — the data could be read straight off the page.
| Site | Listed Staff | Notable Details |
|---|---|---|
| gitnux.org | Diana Reeves and 2 others | The No. 1 product doesn't even appear on worldmetrics.org's list; labeled "AI-verified · Expert reviewed" |
| worldmetrics.org | Kathryn Blake and 2 others | Shows only its own ranking |
| wifitalents.com | Ryan Gallagher and 2 others | Shows only its own ranking |
That's nine different named staffers attached to a single question. On top of that, the byline on two of the three pages literally read "Within the next 26 days," and the third read "Within the next 40 days" — what look like template placeholders that never got rendered.
Vendor Links Were Inconsistent Too
The researchers checked all 1,502 vendor homepages that Perplexity supplied, testing each one via a direct connection and via rotating proxies. Ten addresses didn't resolve at all (eight of those didn't even have a nameserver), and four returned no response — 17 domains in total (1.1%) were unreachable. graphiql.com, presented as GraphiQL's homepage; todo.com, presented as Microsoft To Do's homepage; and aquasecurity.io, presented as Trivy's homepage — none of those sites actually exist. Another 92 (6.1%) redirected to a different registered domain, though most of those were ordinary acquisitions or rebrands.
The remaining two cases are the real problem.
| Category | sonar's answer | sonar-pro's answer |
|---|---|---|
| Research data management platforms | dryad.co → redirects to an Indonesian online gambling site | datadryad.org (normal) |
| Data quality tools | montecarlodata.com (normal) | montecarlo.com → redirects to a Monaco casino group |
In practice, the two models look less like two separate systems and more like the same search engine measured twice. Their citation lists were identical in 289 of the 380 categories, and the Jaccard similarity of their URL sets reached 0.898. The top recommendation matched in 290 of the 380 categories as well.
What the Study Didn't Establish
The researchers were upfront about the study's limits. This measurement covers Perplexity only — ChatGPT, Gemini, Copilot, and Google's AI Mode weren't examined. Because it's a single-day snapshot, grounding results can shift over time, and each prompt was only asked once, phrased one way. The 380 categories were a list the researchers put together themselves, not a sample of questions real buyers actually ask. They also stressed that Tranco rank measures popularity, not quality. Most importantly, they didn't test whether recommendations would change if these sources were removed, and they don't rule out the possibility that the three sites actually recommended reasonable products. As for the shared operator behind the three brands, they said that's an inference from infrastructure clues alone — they couldn't confirm who actually runs them.
Editor's Take
What this study really shows is that the battlefield for search engine optimization has shifted. Content farms used to chase Google's search rankings; the sites uncovered here never expected human visitors at all — they built pages purely to catch an AI's grounding step, going so far as to label their homepages "Facts & Grounding Page." As more AI search products follow Perplexity's model of searching the live web to build answers, more sites built to slip into that citation pool seem likely to follow.
Viewed generationally, AI search is standing roughly where Google's search stood in its early days. It took Google years in the early 2000s to build filters that weighed link quantity and quality and pushed content farms down its rankings; today's AI grounding layer still doesn't have comparable machinery for filtering domain credibility. The numbers back that up — the median first-archive year for the unverified domains cited here is 2020, and 16.6% of them didn't even exist before 2025. In short: if a page turns up in search, it can become a citation, with no credibility check in between.
On the practical side, this study is worth reading both for people who lean on AI search results to pick software and for marketing teams trying to get their own products surfaced by AI recommendations. Buyers would do well to click through and check any vendor homepage an AI gives them, at least once — this study found actual cases where domain redirects sent people to random gambling and casino sites. For marketers, there's nothing wrong with getting your content picked up during AI grounding, but it's worth asking whether flooding markets you don't even compete in with mass-produced list content is a strategy with any staying power.
In the coming weeks, it's plausible that Perplexity or other AI search providers will adjust their domain-credibility filters in response to this report. Trellner has released its full dataset and code under a CC BY 4.0 license, so it looks likely that other researchers or journalists will run the same method on ChatGPT, Gemini, and Copilot next.





Comments