Can You Trust AI Visibility Scores?

SEO SystemsBy Amir Mousavi

I went looking for the paper behind the "up to 40%" figure because I kept meeting the figure without the paper. Finding it was easy. Reading it was the part that changed my mind, and not in the direction people usually point it.

Somewhere between 2024 and now, "AI visibility" became a product category. A dozen tools will sell you a score: how often ChatGPT mentions your brand, what share of Perplexity answers cite your domain, whether Gemini names you in your category. The dashboards look like rank trackers, because familiar things get budget.

So I spent part of 2026 reading the papers underneath those dashboards. There is a gap between what the research measures and what the category sells, and it is wider than I expected. The most-cited published result in this field concerns content that has already been retrieved. It is marketed as a way to become retrieved.

None of that means AI visibility does not matter. It means most of what is sold as a measurement of it is one sample of a probabilistic system, reported to two significant figures.

Quick answer

No, not as point estimates. AI visibility scores are single-run samples of a non-deterministic system, and a July 2026 survey of 45 studies found no technique with a stable, cross-platform causal effect on organic discoverability. What I do instead: run each prompt at least 30 times per surface, which is a floor below every published sampling threshold, report a confidence interval rather than a point estimate, and anchor to first-party citation data where an engine publishes it.

Where does the "40% improvement" figure come from?

It comes from a real, peer-reviewed paper, and it is almost always misread. "GEO: Generative Engine Optimization," by Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan and Ameet Deshpande, reached arXiv on November 16, 2023, was revised June 28, 2024, and was accepted to KDD 2024. The abstract states: "we demonstrate that GEO can boost visibility by up to 40% in generative engine responses." The same paper introduced GEO-bench and warned that "the efficacy of these strategies varies across domains."

Then I got to the method. The experiment modifies source documents already inside the generative engine's context window, then measures how prominently they appear in the answer. Already inside. It studies what happens after retrieval, not whether you get retrieved.

I read that section twice, because an entire product category is sold on the other reading. Selling the first result as though it were the second is this category's founding error, and every claim stacked on top of it inherits the error.

Why do two AI visibility tools give different numbers?

Because they sample, and what they sample is non-deterministic. Ronald Sielinski's "Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement" (arXiv, March 9, 2026, revised June 9, 2026) opens with the reason: "identical queries submitted at different times can produce different responses and cite different sources." He sampled Perplexity Search, OpenAI SearchGPT and Google Gemini across three consumer product topics, daily for nine days plus a run at ten-minute intervals over four hours.

This was the second paper I found, and it is the one that changed how I read a dashboard. Citation distributions "follow a power-law form and exhibit substantial variability across repeated samples," and bootstrap confidence intervals showed that "many apparent differences between domains fall within the noise floor of the measurement process." I had assumed the instability would sit at the top of the list, where a few large domains trade places. It does not. Rankings were unstable "throughout the frequently cited domain set."

So when two vendors hand you two different numbers, you have not caught one of them lying. You have two draws from a heavy-tailed distribution. Run the same tool twice and you get the same disagreement for free.

What does the research actually support?

I found the survey last, and it is the most useful document I read this year. Olivier Martinez's "Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026)" is an arXiv preprint submitted July 15, 2026. It has not been peer reviewed. I am not going to bury that in an article about evidence standards; I lean on it anyway because its claims are checkable against the corpus it lists, which makes it a paper you can argue with. It reviews 45 studies from a November 2023 to July 2026 window and finds that "no reviewed technique shows a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behavior." On the KDD paper specifically, its "widely cited gains are valid within its experimental setting but conditional on a source already being present in a fixed context; they establish neither organic discoverability nor durable traffic effects."

Two of its findings stayed with me, for opposite reasons. The comfortable one: "topical relevance and context position are the most reproducible levers," while "generic heuristics transfer poorly" and "competition can erode individual gains." Relevance and position. That is close to what I would have guessed, and agreeing with a paper is not evidence, but I will take the alignment.

The other finding I did not enjoy at all: "citation-oriented rewrites can impair retrieval." Sit with that for a second. A large share of what this industry sells is citation-oriented rewriting, so its flagship product may be working against the retrieval step you had to clear first. Martinez also finds that commercial audits show "low source overlap, substantial run-to-run variability, and persistent fidelity gaps."

Then there are the numbers everyone quotes about schema markup and AI citation rates. I tried to trace several of them. Every trail ended in an article citing an article, with no published methodology at the bottom, so I am not restating them here. I would rather leave a hole in this piece than launder someone's unsourced statistic through my own site.

Philipp Götza named the problem in Search Engine Land on January 19, 2026: a ladder of misinference running statement, fact, data, evidence, proof, where "to accept something as proof, it needs to climb the rungs of the ladder." Most GEO claims I meet sit on the first or second rung. A few never went near the ladder.

Which tactics does Google's own documentation contradict?

Google's guide on optimizing for generative AI features, updated July 10, 2026, is blunter than I expected a Google document to be. It forecloses three things this category sells. On schema: "Structured data isn't required for generative AI search, and there's no special schema.org markup you need to add." On AI-specific files: "You don't need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search (including its generative AI capabilities), as Google Search itself doesn't use them." On chunking: "There's no requirement to break your content into tiny pieces for AI to better understand it."

What Google does call a requirement is unglamorous: "a page must be indexed and eligible to be shown in Google Search with a snippet, fulfilling the Search technical requirements." Its companion page on AI features, updated December 10, 2025, adds that "there are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary." Google names retrieval-augmented generation over its core Search index as the mechanism, which means the retrieval layer is the index you have been working on for years — the SEO architecture checklist, not a new file format. A boring conclusion, and I think it is the right one.

ChatGPT, Perplexity and Claude document their retrieval far less. That asymmetry limits what anyone can honestly claim about them, including me.

A measurement protocol you can defend

Here is what I would actually run. It is built for one moment: someone asks how you know, and the answer is a method instead of a vendor logo. It also decays the way an event specification decays when nobody owns it, which is the pattern behind why analytics implementations fail and the analytics implementation checklist.

  1. Freeze a prompt set. Twenty to forty prompts a real buyer would type, in a versioned file: category, comparison and problem-first questions. Adding prompts mid-quarter makes trend lines meaningless, and mid-quarter is exactly when somebody will want to add prompts.
  2. Fix the surfaces. Name the assistants and modes, plus account state: logged out, memory off, personalization off, one locale. Different surfaces are different populations. Never pool them, and if a dashboard is quietly summing them for you, it is wrong.
  3. Set repetitions before you see any result. This is the step that gets cut first, usually on cost, and cutting it turns the whole exercise into decoration. For a 95% confidence interval spanning five percentage points on citation share, Sielinski reports roughly n ≈ 40–50 for Gemini, about n ≈ 100 for Perplexity and n ≥ 150 for SearchGPT. Thirty runs per prompt per surface is an absolute floor that clears none of those thresholds; if that is too costly, cut prompts rather than repetitions.
  4. Record run-level rows, not aggregates. Timestamp, surface, prompt ID, whether your domain was cited, its position, and every cited domain. That last column turns a count into a share, and most vendor exports omit it. Ask to see one before you sign anything.
  5. Report an interval, never a point. Bootstrap the runs: "cited in 18% to 31% of runs, n = 30, week of July 20, 2026." (Those figures are a format example, not a measurement.) A single number implies precision the measurement does not have, and the single number is the one people remember.
  6. Re-baseline on a fixed cadence and pre-register changes. Same prompts, surfaces and repetitions each month or quarter. Before a content change, write what you expect to move and by how much. This is the other step everyone skips, and it is the cheapest one here. A predicted movement is evidence; one noticed afterward is a story.
  7. Anchor to a first-party denominator. Bing Webmaster Tools added Citation Share on June 16, 2026: your citations as a share of all citations shown for the same grounding query. Its AI Performance preview (February 10, 2026) adds Total Citations and Average Cited Pages. The full definition and the other three measurement layers are in how to measure AI search traffic.

Claim, evidence, action

ClaimWhat the evidence supportsWhat to do instead
"GEO lifts visibility up to 40%."True for already-retrieved sources in a fixed context (Aggarwal et al., 2024).Measure "am I retrieved?" and "how am I presented?" apart.
"Add this schema and get cited more."Google, July 2026: no special schema.org markup is needed.Keep schema for what it does do.
"Rewrite pages for citation."Martinez, 2026: "citation-oriented rewrites can impair retrieval."Work on topical relevance and context position.
"Your score is 42."Single runs sample a power-law distribution with high variance (Sielinski, 2026).Report a bootstrap interval with n and date.
"Ship an llms.txt to reach AI answers."Google: AI text files are not needed for Google Search.Ship one if you want; do not book it as a tactic.

Does llms.txt do anything? What this site claims

This site ships an /llms.txt. I keep it because it is a cheap, honest table of contents describing my pages in my own words. Whether assistants read it, I do not know. I have no evidence it changes whether any of them cite me, and Google's July 2026 guidance says such files are not needed. It costs one route file, so I keep it and I do not claim it works. Where I use AI in content work, the useful parts are research, briefs, metadata and QA rather than citation-chasing rewrites — the workflow in AI agents for SEO content.

My working take

I am not against these tools. Sampling assistants at scale is real engineering, and I would rather someone else ran that infrastructure. What I object to is the number on the front of the dashboard, which is one draw from a high-variance distribution, printed like a fact.

If you are about to buy one, the useful questions are not about features. How many runs per prompt, per surface? Can I see the prompt set, edit it, and keep its version history? Does the export list every cited domain or only mine? A vendor who answers those is doing measurement. A vendor who answers with a score out of 100 has built a rank tracker for a system that does not rank.

The honest version of my position is less satisfying than a score. I do not know whether most GEO tactics work. Neither does anyone selling them, and that is the 2026 literature's finding, not my opinion about it. What I do know is that the retrieval layer is the index, that improving it looks a great deal like the SEO work that already existed, and that a measurement you can defend in a meeting beats a number you cannot.

Frequently asked questions about AI visibility scores

Can you trust AI visibility scores from commercial tools?

Not as point estimates, no. Martinez's July 2026 survey reports that commercial audits show "low source overlap" and "substantial run-to-run variability." Ask a vendor for its sample size and prompt set, then compare intervals rather than numbers.

How many times should you run a prompt to measure AI visibility?

More times than you will want to. At least 30 runs per prompt per surface, and that floor clears none of the published thresholds. For a 95% confidence interval spanning five percentage points on citation share, Sielinski (2026) reports roughly n ≈ 40–50 for Gemini, about n ≈ 100 for Perplexity and n ≥ 150 for SearchGPT. If the budget will not stretch that far, measure fewer prompts properly instead of more prompts badly.

Does the "40% GEO improvement" apply to my website?

Only in the narrow sense the paper tested. Aggarwal et al. (KDD 2024) measured how sources already in a generative engine's context can be modified to appear more prominently, and Martinez calls those gains "conditional on a source already being present in a fixed context." If you are not retrieved in the first place, the finding has nothing to say about you.

Does structured data improve AI citations?

Google's guidance updated July 10, 2026 says no special markup is required: "there's no special schema.org markup you need to add." I still ship schema, for the things schema actually does. The vendor statistics claiming large schema-driven citation lifts are exactly the ones I could not trace to a methodology.

What is the most reliable AI visibility data available today?

First-party engine data, and it is not close. Bing Webmaster Tools has reported Citation Share since June 16, 2026, calculated against all citations shown for the same grounding query. It is a denominator you did not have to estimate.

Sources I used

Verified at source: July 29, 2026.