"AI visibility score" sounds like one number. It isn't. Across the category it refers to at least three different quantities: whether your brand was named at all, how prominently it appeared, and whether your URL made the source list. Two tools can watch the same brand on the same day across the same engines and report different numbers, and both can be correct, because they are measuring different things. Underneath all of them sits a harder problem: none of these tools can see inside the engines. They run prompts and observe outputs, which makes every score a sample of a stochastic process rather than a reading off a meter. This is a guide to what that buys you, what it doesn't, and what to ask before signing.
Start here: everything is a sample
No AI visibility vendor has access to OpenAI's, Google's or Anthropic's internal retrieval logs. What they have is an API key and a list of prompts.
The method is essentially the same everywhere. Build a prompt set meant to represent how buyers ask about your category. Run those prompts across several engines on a schedule. Parse the responses for brand names and cited URLs. Aggregate into a percentage. Charge monthly.
That's a reasonable approach to an otherwise invisible problem, and the category exists because the alternative is knowing nothing. But <u>the score you receive is an estimate produced by sampling, and every methodological choice in that pipeline moves the number.</u> Most of those choices are made by the vendor and not shown to you.
Google, for its part, warns against third-party tools that claim to use internal Google metrics, noting flatly that no third-party tool has access to its ranking or AI systems. That's worth holding onto when a dashboard presents a score to two decimal places.
The three things "visibility" means
When a vendor says visibility, ask which of these they mean.
Mention rate. Did the brand name appear anywhere in the generated answer? Simple to compute, and the most forgiving definition. A passing reference in a list of eight alternatives counts the same as a direct recommendation.
Position-weighted share. Where did the brand appear, and how much answer real estate did it occupy? Being named first in a three-item recommendation is scored higher than a mention in a closing aside. This is closer to what marketers actually care about and harder to compute consistently, because it requires the vendor to define a weighting scheme that they rarely publish.
Citation share. Did one of your URLs appear in the response's source list? This is a different question from whether your brand was named, and the two diverge constantly. A model can recommend your product using knowledge from its training data while citing a competitor's comparison page as the source. You'd score well on mention rate and zero on citation share.
There's a fourth distinction that matters more than any of the above and almost never appears on a dashboard: the difference between making the source list and actually shaping the text. A page can be cited in a footnote while contributing nothing to what the answer says, and a page can shape two paragraphs while being cited once. Appearing in the sources is not the same as being the source, and no widely available tool cleanly separates the two.
None of this is a criticism of the vendors. These are genuinely hard measurement problems. It is a criticism of buying a single number without knowing which one it is.
Four dials that move your score without your content changing
Before you attribute a score change to your content work, check whether any of these moved.
Prompt set composition. You or the vendor choose the prompts, which means you choose the denominator. A prompt set weighted toward branded queries produces a high score. The same brand measured on category and comparison prompts produces a much lower one. Neither is wrong; they answer different questions. Adding or removing a handful of prompts can shift a score by more than a quarter of good content work would.
Competitor set. Share of voice is your mentions over total mentions across a tracked brand set. Narrow that set and your share rises with no change in reality. Widen it and your share falls. If the competitive set isn't frozen across periods, the trend line is meaningless.
Samples per prompt. Run a prompt once and you get one draw from a distribution. Run it ten times and you get a rate. Vendors differ substantially here, and the difference between one run and ten is the difference between noise and signal. This is the single best question to ask a salesperson.
Engine mix. The same brand and the same prompts produce very different results across engines, because they use different indexes and different retrieval. A blended cross-engine score hides that. If the vendor reports one number, ask how it's weighted.
Three of those four dials are set by whoever configured the tool, which means the score is partly a description of your configuration rather than of your market.
The non-determinism problem nobody puts on the pricing page
Here's the part that should change how you read every AI visibility chart you've ever seen.
Language models do not return the same answer to the same question. That's expected at normal settings, since generation samples from a probability distribution. What surprises most marketers is that it remains true at temperature zero, the setting that's supposed to make output deterministic.
The cause is infrastructural rather than statistical. As Thinking Machines Lab set out in its analysis of the problem, reproducibility is remarkably difficult to get out of these systems even under greedy decoding. Cloud APIs route requests across different hardware, batch concurrent requests in varying combinations, and floating-point arithmetic isn't associative, so summation order changes intermediate values. Small numerical differences ripple into different token choices. Research on multi-turn model behaviour notes the same thing and points out that the major providers acknowledge it: Anthropic recommends sampling multiple times to cross-validate consistency, OpenAI offers a seed parameter as a best-effort improvement rather than a guarantee, and Google describes its outputs as mostly deterministic. Work reported by Jalil and colleagues found ChatGPT returning non-deterministic answers to simple prompts roughly 10% of the time even at temperature zero.
Now the finding that matters most for brand measurement. A 2026 study taking the first systematic look at non-determinism at the token-probability level found the effect is negligible when a token's probability sits near 0 or 1, and significant when it sits between roughly 0.2 and 0.8.
Think about what occupies that middle band. Not grammar, not obvious facts. <u>It's exactly the decision of which brand to name when several are plausible, which is the decision your entire visibility score is built on.</u> Non-determinism is smallest where the model is confident and largest precisely where your competitive position is contested.
Two practical consequences. First, a single measurement run tells you close to nothing about a contested prompt, and month-over-month movement on a small prompt set is often variance rather than performance. Second, brands with entrenched dominance will measure stably, while everyone competing in the middle will see noisy scores. If your category is genuinely competitive, expect your chart to be jumpy for reasons that have nothing to do with your content.
Ask any vendor how many runs per prompt per period they execute, and whether they report variance alongside the mean. A tool that reports a point estimate with no error band is presenting a sample as a certainty.
What these tools cannot do: connect citations to revenue
This is the honest limit of the entire category, and it isn't a software problem that a better product will solve next quarter. It's a data problem with three disconnected systems and no join key.
System one: the citation data. Your visibility tool knows your brand appeared in an answer to a synthetic prompt it ran itself. It does not know whether a real person ever asked that question, or how many did.
System two: the traffic data. GA4 added a native AI Assistant channel on 13 May 2026, automatically assigning an ai-assistant medium to sessions arriving from recognised AI referrers, with wider availability following in June. It's a real improvement over maintaining regex by hand. It's also incomplete in ways that matter. Traffic arriving without a referrer header, which includes in-app browsers, mobile apps and copy-pasted links, still lands in Direct and is indistinguishable from someone typing your URL. Google hasn't published the full list of recognised referrers. And clicks from Google's own AI Overviews and AI Mode are counted as Organic Search rather than as AI traffic at all.
System three: the search data. Search Console's Search Generative AI performance reports, launched 3 June 2026, show impressions inside AI Overviews and AI Mode broken out by page, country, device and date. They report impressions only. No clicks, no CTR, no query data.
Three measurement systems, none of which shares an identifier with the others. The citation tool observes prompts nobody asked. GA4 observes a subset of clicks with the sources partly obscured. Search Console observes impressions with no clicks attached. Any claim to trace a citation through to a closed deal is bridging those gaps with an assumption, and the assumption is doing all the work.
Correlation across the three is still useful. Citation share rising while AI-channel sessions and assisted conversions rise is meaningful evidence, and it's the strongest evidence currently available. It is not attribution, and a vendor who calls it attribution is telling you something about their sales process rather than their product.
The influence you will never measure
One more limit, which applies to every tool including any built in future.
A person asks an assistant which vendor to shortlist. Your brand is named. They never click, because they got what they needed. Two weeks later they search your brand directly and convert, and your analytics records a branded organic session.
The AI answer did the persuading and the branded search took the credit. No referrer exists to capture, no click happened to attribute, and the visibility tool that observed a similar prompt has no way to link its observation to that person. As zero-click behaviour grows, the share of AI influence that is structurally invisible grows with it. Any model claiming to quantify total AI-driven revenue is estimating this, not measuring it. Ask to see the estimation method.
Twelve questions before you sign
Take these to the demo. The answers will separate the serious vendors from the dashboard-with-a-markup.
- Which definition does your visibility score use: mention, position-weighted, or citation?
- How many times do you run each prompt per measurement period?
- Do you report variance or confidence intervals, or only a mean?
- Who controls the prompt set, and can I freeze it across periods?
- Who controls the competitive set, and can I freeze that too?
- Which engines are included, and how is a blended score weighted?
- How do you handle a model version change mid-period? Is it annotated?
- Do you distinguish being cited in the source list from influencing the answer text?
- Do you measure sentiment, and how is it validated against human judgment?
- What exactly do you claim about revenue attribution, in writing?
- Can I export raw responses, or only aggregates?
- What geography and personalisation settings do your prompts run under?
Question three is the one most likely to produce a pause. A tool that cannot tell you its own error bars is asking you to make budget decisions on a number whose precision it hasn't measured.
How to actually use the number
None of this means don't buy one. It means calibrate what you expect.
Treat the score as a relative trend on a frozen configuration. Lock the prompt set, lock the competitor set, lock the engine list, and change them only at deliberate, annotated intervals. The absolute number is close to meaningless as a benchmark against another company using a different configuration. The direction of travel on your own frozen setup is genuinely informative.
Report per engine rather than blended, because the strategic implications differ and the blend hides them. Annotate model releases, since a version change can move your score more than a quarter of content work. And set the reporting cadence to monthly rather than weekly for anything but the largest prompt sets, because at small sample sizes a weekly chart is mostly rendering noise as narrative.
Then pair it with the click-side data you do control, accepting the seams. Rising citation share alongside rising AI-channel sessions and stable post-click behaviour is a good quarter. Rising citation share with flat sessions might mean your citations aren't in answer positions that produce clicks, or might mean the traffic is landing in Direct where you can't see it. Sitting with that ambiguity honestly is better than buying a tool that resolves it with a confident number it cannot support.
FAQ
Do different AI visibility tools give different numbers for the same brand? Yes, and they should, because they use different prompt sets, competitive sets, sampling rates, engine coverage and scoring definitions. Differing numbers are evidence of methodological difference, not of one tool being broken.
Why does my AI visibility score change when my content hasn't? Four common causes: the prompt set or competitor set changed, the vendor's sampling rate changed, an engine shipped a model update, or you're seeing run-to-run variance. Language models return different answers to identical prompts even at temperature zero, and that variance is largest for genuinely contested choices such as which brand to recommend.
Can any tool prove AI citations drove revenue? No. Citation data, GA4 traffic data and Search Console data are three separate systems with no shared identifier, and a large share of AI-influenced conversions arrive with no referrer or as branded search. Correlation across those systems is achievable and useful. Causal attribution is not currently available from anyone.
Does GA4's AI Assistant channel solve AI traffic measurement? It solves part of it. Launched 13 May 2026, it auto-classifies sessions from recognised AI referrers without setup. It doesn't capture traffic arriving without a referrer, which still lands in Direct, and Google's own AI Overviews and AI Mode clicks are counted as Organic Search.
Can Search Console show me AI Overview clicks? No. The Search Generative AI performance reports launched in June 2026 report impressions, pages, countries, devices and dates. Clicks, CTR and query data aren't included.
How many prompts do I need for a reliable measurement? There's no evidence-backed universal number, and anyone quoting one precisely is guessing. What matters more than prompt count is runs per prompt, since a single run on a contested prompt is one draw from a distribution. Ask vendors for their sampling design and whether they publish variance.
Is sentiment analysis in these tools trustworthy? Treat it as the weakest metric in the stack. It's usually an LLM classifying another LLM's output, which stacks two probabilistic systems, and validation against human raters is rarely published. Useful for spotting large directional shifts, not for small deltas.
Should I just build this in-house? You can run prompts manually against several engines and log results in a spreadsheet, and for a small prompt set that's a legitimate baseline. It stops scaling quickly once you want multiple runs per prompt across several engines and competitors. The build-versus-buy line is usually about sampling volume, not about capability.
Where RankSage fits
The gap this post describes is a join problem. Citation data sits in one system, click data in another, and post-click behaviour in a third, which is why "did our AI visibility work pay off" is currently answered with a shrug or with a number nobody can defend.
RankSage is being built to close part of that gap: citation tracking across ChatGPT, Claude, Gemini, Perplexity and Copilot, joined per page with GA4 and Search Console data and first-party behavioural signals, so citation movement can be read against what actually happened on the page.
Worth saying plainly, since the whole point of this piece is honesty about limits: that join does not fix the underlying problems. RankSage will sample the same non-deterministic engines everyone else samples, and it will not see AI traffic that arrives with no referrer, because nobody can. What joining the data does is let you correlate movements that are currently sitting in separate tools, and make the seams visible instead of papering over them. It isn't launched yet. If you're evaluating this category now, join the waitlist for early access.
