A colleague sent me a screenshot last month. An AI answer had cited one of our pages to support a claim that page had never come close to making. Not a paraphrase issue. Not a slight stretch. The page was about topic A. The answer used it as evidence for topic B. They shared a topical neighborhood but the page literally contradicted the claim it was being used to back up.
That one screenshot kicked off something I should have done months ago. I went through every AI citation to our domain I could find across ChatGPT, Perplexity, and Gemini over a 60-day window. 80 citations total. For each one I opened the cited page, read the surrounding context, and asked one question: does this page actually support the specific claim the answer attributes to it?
50 of the 80 citations were fine. The page said roughly what the answer claimed it said. Maybe loose paraphrasing here and there, occasional oversimplification, but nothing that would make me uncomfortable. Standard extraction behavior.
7 citations had issues. Not huge, but noticeable. The answer pulled a true statement from the page but framed it as supporting a different point than the original author intended. Sort of like quoting someone out of context except there's no malice, just a model matching keyword overlap to semantic proximity and occasionally missing the mark. Annoying but livable.
Then there were the 23 that genuinely worried me. These weren't paraphrase stretches or context shifts. The answer made a specific factual claim, attached our URL as the source, and our page did not contain that claim. In some cases our page said the opposite. In others the page had simply never addressed that question at all. The model seemed to be citing us based on topical relevance rather than factual support. Close enough in subject matter that the URL looked plausible as a source, wrong enough that anyone who actually clicked would realize the citation was bogus.
What bothers me about this isn't the error rate. 23 out of 80 is 29 percent, which sounds bad until you consider that I was specifically hunting for problems and may have selection-biased the sample toward ambiguous cases. The real number could be lower. Could also be higher if I checked more systematically.
What bothers me is that nobody in GEO seems to be tracking this. We obsess over citation counts. We build strategies around increasing them. We treat every new citation as a win. But if nearly a third of those wins are attributing claims you never made, what exactly are we winning? Brand visibility for wrong ideas? Traffic from people who click through and find irrelevance?
I'm starting to think citation count might need a quality filter we're not measuring yet. Not just "did an AI name-drop our URL" but "did it name-drop us for something we actually said, and something we'd stand behind." Those are different outcomes and the current tooling conflates them completely.
And there's a trajectory problem. As AI answers get more confident-sounding and citations become smaller footnotes that fewer users verify, the incentive for accuracy on the model side might actually decrease. The citation becomes a trust signal for the answer rather than a factual anchor. And if that's the direction we're heading, being highly citable starts to look different than I thought it did. You want to be cited for the right things, not just cited often.
Source: r/GEO_optimization · by /u/Brave_Acanthaceae863