A researcher in a patterned jersey working late at a monitor-lit desk covered in sticky notes and printouts, documentary photography
News

215,128 Machine-Written 'Best Software' Pages Are Feeding AI Answers

Three sites with a shared template published 215,128 buying guides and describe themselves as 'Facts & Grounding Pages'. Perplexity's search models cite them more often than Gartner.

PerplexityAI SearchSEOCitationsAI Hallucination

The clearest evidence yet that AI search engines are being farmed didn’t come from a watchdog or a rival. It came from someone doing the boring thing: asking 380 ordinary buying questions, writing down every link the engine retrieved, and looking the domains up.

The findings, published September 2 by Trellner, are blunt. Nearly 60 per cent of the 7,534 citations returned by Perplexity’s sonar models across 380 software categories pointed at domains ranked worse than 100,000th in the web’s traffic rankings. 23 per cent pointed at domains not in the top million at all. Wikipedia — the internet’s most-cited reference — was cited three times.

The machine-written buying guide factory

The headline find sits in the top ten sources. Three sites — wifitalents.com, worldmetrics.org and gitnux.org — appear to be one operation. Same NameCheap registrations between December 2023 and May 2024, same Cloudflare nameservers, same page template, six blog posts each, and sitemaps listing roughly 100,000 URLs apiece. Between them: 215,128 machine-generated “best [category] software” pages. There are not 215,128 software categories. The report’s authors note all three are registered through the same infrastructure; they’re careful to call that strong circumstantial evidence of common control rather than proof of ownership.

But the detail that should follow the AI-search industry around is what two of these sites call themselves. Their homepages carry the HTML title “Facts & Grounding Page” — grounding being the exact term of art for the retrieval step where a search model fetches documents before answering. Their meta descriptions describe “an independent market research company publishing… software Best Lists,” written, per the sites’ own framing, as “one machine-readable record.” No human buyer reads a title tag at 2am looking for a CRM. These pages are addressed to the software that reads them. The report is careful with its language — it stops short of accusing anyone of deception, noting the sites publish “in the ordinary way” and that the measurement is about what retrieval layers do with such content.

The vendor blog that outranks Gartner

Second finding, less lurid and more damning: the third-most-cited source overall is guideflow.com — a company that sells interactive product demos and competes in none of the 380 categories tested. Its content-marketing blog, 96 distinct URLs, was cited 194 times, ahead of Gartner, supplying the evidence base for questions about 3D rendering software, IVR software and architecture software alike. Nothing deceptive, as the report says. Just a vendor’s own listicles about markets it doesn’t operate in becoming the third-largest authority in an AI answer engine’s evidence base.

Perplexity did not respond in the report, and the authors scope their claim narrowly — only Perplexity’s models were measured, only through OpenRouter, and they explicitly say nothing should be read about other engines. That care matters. An earlier audit this week found a third of Perplexity’s figure citations didn’t support the sentence they footnoted; this one is upstream of that problem. It’s not that the citations break. It’s that the pool they’re drawn from is increasingly manufactured. The BBC’s earlier investigation into AI citation poisoning described the same mechanism at Google; this audit adds scale and a method anyone with a spreadsheet can repeat.

Why this is worse than SEO spam

Junk pages ranking on Google were an annoyance you scrolled past. Junk pages feeding AI answers are structurally different: the reader never sees the ten blue results, only the confident summary and its footnotes. If the retrieval layer can be seeded — and 215,128 template pages is seeding at industrial scale — then whoever manufactures the pages writes the answer. The sites in question aren’t hiding it, either. They describe themselves, in machine-readable form, as exactly what they are.

What stands out to me is the honesty asymmetry. The researchers published their full method — 760 calls, every domain looked up, every vendor homepage fetched, limitations stated. The sites exploiting the pipeline published a compliance page optimised for the crawler. The web’s incentive gradient now points at machines, not readers, and the machines can’t tell the difference.

The New Zealand angle

Small businesses buying software off AI recommendations should treat the footnotes the same way we’ve been saying all week: click them. If the “best CRM” list traces back to a site you’ve never heard of, registered last year, with 100,000 near-identical pages — you’ve found the answer’s manufacturer, not its evidence.

FAQ

What are “Facts & Grounding Pages”? HTML pages whose titles and meta descriptions are addressed to AI retrieval systems rather than human readers, describing a site as a machine-readable record of verified facts.

How many machine-generated pages did the audit find? 215,128 “best software” buying guides across three domains with a shared template, none of which existed before December 2023.

Did the audit cover other AI search engines? No. Only Perplexity’s sonar and sonar-pro models, tested through OpenRouter. The authors explicitly say nothing should be inferred about other engines.

Sources: Trellner, Hacker News