Here is a test anyone can run: an AI search engine gives you an answer with a number in it, a footnote next to the number, and a link. You click the link. Does the page actually contain the number?
Haus Research ran that test 1,826 times against Perplexity’s search models — 310 factual questions about 210 technology companies, every cited URL fetched and checked. The results, published on Tuesday: 34.7 per cent of citations attached to a figure pointed at a page that either wouldn’t open, or opened and contained none of the numbers in the sentence it was cited for.
The detail that makes this worse than a routine accuracy complaint: it isn’t mainly hallucination. Only 1.3 per cent of the cited URLs were dead. The failure is structural. 16.1 per cent of cited pages sit behind paywalls, logins, or bot walls — a footnote a reader cannot open is not proof, whatever it says. The rest are live pages that simply don’t say the thing.
Scored generously, it still fails
To their credit, the auditors stress-tested their own method. They scored per claim rather than per citation — counting a claim as passing if any one of its several citations carried one of its figures — and 14.4 per cent still failed. They retried blocked URLs through a rotating proxy so pages wouldn’t be written off over one unwelcome datacentre address. A bare year in the cited page counted as a pass. The 34.7 per cent is a floor, and the authors are explicit that it’s a floor.
There’s also a subtler finding that deserves more attention than it will get: a quarter of the cited pages — 25.1 per cent of the 1,500 checked — have never been captured by the Wayback Machine at all. Not once. The citation layer under AI answers is being assembled out of pages built to rank, not to last, and when one disappears the sentence and the footnote stay while the evidence evaporates. We’ve written before about AI-generated junk pages being manufactured to feed AI citations; this audit measures how dependent the citation layer has become on exactly that kind of page. Directory sites built to rank — PitchBook, Tracxn, LinkedIn pages — take a quarter of the sourcing, and nearly two in five of those have no archived copy.
The part that should sting Perplexity most
Perplexity’s whole pitch is that it shows its work. The company is already suing news publishers in one direction — CNN sued Perplexity over 17,000 articles — while its agentic computer-use plans push the same answer-engine deeper into daily workflows. An audit showing the footnotes don’t function as proof undercuts the specific thing customers are paying for. It has also landed days after Perplexity’s CEO argued AI job losses are “worth it” — a claim whose evidentiary base now looks shakier.
Some fairness is due: the audit tested Perplexity’s models through OpenRouter, not the consumer product, and the authors list their own limitations — English only, tech companies only, HTML fetched without JavaScript. A gated page isn’t necessarily a wrong page. But the question they measured is the right one: not “is Perplexity accurate?” but “does the thing it offers as proof function as proof?” On a third of the citations, it doesn’t.
For New Zealand readers, the practical note is the one we keep returning to: AI-cited sources have failed audits before, and academic publishers are rejecting papers over fabricated references. An inline citation from an AI search tool is a claim about provenance, not provenance itself. Click it. Especially before you repeat the number.
FAQ
How many Perplexity citations failed the audit? 34.7 per cent of the 1,826 citations attached to sentences containing figures pointed at pages that either couldn’t be opened or didn’t contain any of the numbers in the sentence. Scored per claim, 14.4 per cent failed.
Were the failures hallucinated links? Mostly no. Only 1.3 per cent of cited URLs were dead. The largest failure modes were paywalled or bot-blocked pages (16.1 per cent) and live pages that don’t contain the cited figure.
Does this apply to the Perplexity app I use? Not directly — the audit tested Perplexity’s sonar models via OpenRouter, not the consumer product. But the underlying retrieval-and-cite mechanism is the same, and no independent audit of the consumer product’s citation accuracy has been published.
— CJ Murden, editor of Singularity.Kiwi. Former digital technologies teacher, author of AI-focused books. Writing with a New Zealand focus.