
How is AI answer visibility defined and measured?
Visibility, contribution, stability: three different values stacked under the one word 'visibility'
How is AI answer visibility defined and measured? Three metrics with formulas and paper figures: visibility (GEO), contribution (Shapley), and stability. Saying a brand was 'visible' in an AI answer translates into at least three values: visibility (how much and how early), contribution (did it actually shape the answer), and stability (does it hold up when repeated).
Translate "our brand was visible in an AI answer" into formulas and you get at least three different values: how much and how early it appeared in the answer (visibility), whether that source actually shaped the answer (contribution), and whether a single measurement can be trusted (stability). The three values cannot stand in for one another. Visibility can be high while contribution is low, and both can swing widely on the next measurement if you measure only once. Underneath the one sentence "GEO is working" sit three different questions.
Here is the scope of the evidence. The three definitions come from the original GEO paper [1], the RAG source attribution paper [2], and the Don't Measure Once paper [3], and we checked the formulas and figures directly against the original text of all three. Every figure holds under the paper's experimental conditions. We have a stake here: we run TRAIL Search, which measures AI search visibility. So at the end we also state which of the three definitions our product measures and which it does not.
We have covered each paper on its own. The visibility metrics and nine methods are covered in detail in The paper that defined GEO: visibility metrics, 9 methods, and splitting source contribution with Shapley values in Which Documents Did RAG Use? Shapley Source Attribution. Rather than repeat those summaries, this post focuses on what you see when the three definitions sit side by side.
Three definitions of generative search visibility metrics
Generative search visibility metrics fall into three main definitions: visibility, contribution, and stability. The three definitions measure the same thing in different units. Visibility looks at the share a source takes within a single answer, contribution at the role a document plays while the answer is produced, and stability at how well results hold across many answers. Say "it was visible" based on only one of them, and the other two questions are left unanswered.
| Definition | Question asked | Unit measured | Key paper | What you miss if you look only at this |
|---|---|---|---|---|
| Visibility | How much and how early did it appear in the answer | Length and position of citing sentences within one answer | GEO (KDD 2024) | Whether the source was actually used in the answer |
| Contribution | Did that source actually shape the answer | Change in the answer's likelihood as documents are added and removed | Source Attribution in RAG (2025) | How noticeable it was on the user's screen |
| Stability | Can a single measurement be trusted | Overlap of source sets and standard error across repeated runs | Don't Measure Once (2026) | The size and meaning of the value itself |
Table 1. The three definitions of AI answer visibility, and the question each cannot answer on its own [1][2][3].
Visibility: how much and how early in the answer
Visibility is the definition that measures a source by the length and position it takes up inside the answer. The GEO paper proposed this as an impression metric for generative engines in place of rankings [1]. The simplest form divides the word count of the sentences citing a source by the word count of the whole response.
is the set of sentences citing source , is all sentences in response , and is the number of words in sentence . The original has one more rule: when several sources cite the same sentence, its word count is shared equally among those citations [1].
Word count alone treats a citation the same whether it comes early or late. So the paper also uses a position-adjusted word count that normalizes a sentence's position by the number of sentences in the response and decays it exponentially [1].
is the position of sentence in the response, and in the exponent's denominator is the number of sentences in the response. Earlier citations count for more. The paper grounds this choice in prior studies showing that click-through rates follow a power law with ranking in search results. All impression metrics are normalized so that the impressions of all citations in a response sum to 1, and a method's effect is compared as relative improvement of the modified response over the original [1]. By this measure, top methods such as statistics addition, quotation addition, and citing sources showed a 30–40% relative improvement in position-adjusted word count (under the paper's experimental conditions) [1]. The code and benchmark are public in the GEO project repository [5].
The strength of visibility is that you can compute it from the answer text alone. You don't need access to the model's internals. Its limit comes from the same place. It looks only at the length and position of sentences carrying a citation marker and never asks whether the source actually shaped the content of the answer.
Visibility is also relative. C-SEO Bench measured the gain in citation rank while varying the share of documents adopting the same optimization method, and showed a zero-sum structure in which early adopters' gains shrank toward zero as adoption rose [4]. When one source's visibility rises, other sources in the same answer get a smaller share. The benchmark data is public as a Hugging Face dataset [6]. Detailed results are in C-SEO Bench: Does conversational SEO actually work?.
Contribution: being cited in an AI answer versus actually contributing to it
Contribution is the definition that measures a source by how much the answer relied on it. The Source Attribution in RAG paper treats retrieved documents as players in a cooperative game and computes document 's share as a Shapley value [2].
is the full set of retrieved documents, and is a subset of documents excluding . Document 's share is the weighted average of how much utility rises when is added to every possible subset. Utility is defined as the log-likelihood that the model generates the original full-context answer given only the query and the document subset [2].
The more likely the original answer becomes with a given set of documents, the higher that set's utility. If visibility measures "was it seen," contribution measures "was it relied on."
The definition is theoretically clean, but it does not hold as-is in real models. Shapley values satisfy axioms such as symmetry, which gives equal scores to equally contributing players, and the paper notes that because the log-likelihood utility is not additive, these axioms generally do not hold in the RAG setting [2]. The experiments exposed the gap. With two near-duplicate documents, the one placed first consistently received the higher score, and swapping the order flipped the scores. In multi-step reasoning, where a fact is drawn from one document and the answer completed with another, scores piled onto the document that directly contained the answer, and the document needed for synthesis was undervalued (under the paper's experimental conditions) [2].
Cost is another issue. Every utility evaluation requires a model call, and exact Shapley values require evaluating every subset of documents. That is why the paper compared approximations such as Kernel SHAP and ContextCite, and the two methods that fit a surrogate linear model reproduced the exact Shapley rankings best [2]. And because utility is defined through the model's token probabilities, this value cannot be computed from the answer text alone.
Stability: why AI search visibility shouldn't be measured only once
AI search visibility shouldn't be measured only once because cited sources change a lot even when the same question is repeated [3]. Stability is the definition that measures how well results hold when the same measurement is repeated. The Don't Measure Once paper measured how much the cited-source sets of two observations overlap using Jaccard similarity [3].
and are the sets of sources cited in two runs. A value of 1 means both runs cited the same sources, and 0 means they cited entirely different ones.
The results are far from what search engines have taught us to expect. Running eight prompts per vertical across four Swiss-German verticals through ChatGPT, Gemini, Google AI Mode, and Perplexity, cited-source Jaccard between runs repeated within 24 hours averaged 0.32–0.43 by vertical. That is the same range as the day-to-day figure of 0.34–0.42 [3]. Most of the wobble comes from model randomness rather than change over time. At the brand level, day-to-day Jaccard was 0.45–0.59, higher than at the source level, so whether a brand appears was a more stable metric than individual URLs [3].
Repetition reduces variance. The standard error of a per-brand detection rate was 0.246 at two runs and fell below 0.10, to 0.081, at seven [3]. Citation concentration also differs by engine: the mean Gini coefficient of source citations was 0.715, with Google AI Mode the most concentrated at 0.782 and Perplexity the most evenly spread at 0.671 [3]. That is why the authors recommend setting baselines separately for each engine.
Stability says nothing about the size of the value. Whether visibility is 0.3 or 0.1, or whether contribution is large or small, a stability metric does not answer. Instead it tells you whether a value measured under the other two definitions will survive the next measurement.
Which metrics should be used to judge GEO performance: read the three together
Judge GEO performance with metrics that state which definition they measure, and read the three definitions together. Put the three definitions together and it becomes clear what a single visibility number can hide. A source can have high visibility and low contribution: it appears early in the answer in a long citing sentence, yet removing that document barely changes the answer. The reverse also happens: in multi-step reasoning, a document that supplied the material for the answer may receive only a short citation marker. And either way, if you measure once, more than half of the cited sources may change on the next measurement.
| Situation | Visibility | Contribution | Stability | How to read it |
|---|---|---|---|---|
| Cited early and at length, but the answer is unchanged without it | High | Low | Check separately | Surface share and actual role differ |
| Supplied the answer's material, but got a short citation marker | Low | High | Check separately | Undervalued by visibility metrics alone |
| Came out high in one measurement | High | Unknown | Low | Preliminary observation; measure again |
| Consistently high across many measurements | High | Unknown | High | Visibility is solid, but contribution is a separate question |
Table 2. Typical combinations that appear when the three definitions overlap, and how to read them. The properties of each definition follow the three papers [1][2][3].
The three definitions also require different levels of access. Visibility and stability can be measured from outside with only the answer text and the citation list. Contribution needs the model's token probabilities, so it cannot be computed the same way for commercial engines whose internals you cannot access. That is why the "visibility" in practical reports is mostly a visibility-type value, and the report itself should say so.
TRAIL Search separates what it can measure first
TRAIL Search records the definitions that can be measured from outside first, and keeps them separate. Mentions, citations, and share of voice inside answers are recorded separately for each engine, and a combined value comes only after that. Each cell is labeled by run count as a preliminary estimate (fewer than 3 runs), medium confidence (3 to 6), or high confidence (7 or more), so stability shows up before the value does. These bands are our own rule, set against the standard error figures from the Don't Measure Once paper above [3]. Cells we could not measure stay unmeasured rather than zero.
The second definition, a source's actual contribution, is not part of what TRAIL Search currently measures. The visibility in our reports is a visibility-type value, and the stability label tells you how much to trust it. Why unmeasured and zero should be kept apart is covered separately in How to tell zero from unmeasured in an AI visibility report.
Limitations
Each definition comes from its own experimental conditions. The GEO paper's visibility metrics were validated on a generative engine setup the authors built and on GEO-bench queries, and the paper itself states that the effects of methods may change as engines evolve and that, as a black-box approach, causation is not fully established [1]. The source attribution paper ran its experiments on open models, LLaMA-3.2-8B, Mistral-7B, and Qwen-3B, with BioASQ, Natural Questions, and synthetic scenarios the authors built [2]. There is no basis for assuming these results describe how commercial generative search engines actually behave.
The Don't Measure Once figures come from data collected on servers in Switzerland across four Swiss-German verticals from January to March 2026 [3]. The paper does not say whether Korean prompts or Korean search surfaces would show the same range of variation. C-SEO Bench is likewise limited to English content and two tasks [4].
This post lines up the three definitions side by side and does not judge which one is right. The three values answer different questions. So when a report says "visibility," it needs to say which definition it measured and how many times before anyone can read it.
Frequently asked questions
What definitions of AI answer visibility are there?
Three main ones. Visibility in the answer looks at how much of the answer a source takes up and how early it appears (the impression metrics from the GEO paper). Contribution looks at how much the answer actually relied on that source (Shapley source attribution). Stability looks at how well results hold up when the same prompt is repeated (Jaccard similarity and standard error).
How is being cited in an AI answer different from actually contributing to it?
A citation is the fact that a source is marked on the surface of the answer. Contribution is measured by how much less likely the original answer becomes when you remove that document. In the paper's experiments, of two documents with the same content, the one placed first got the larger contribution score, and documents needed to synthesize the answer were undervalued.
Why shouldn't AI search visibility be measured only once?
Because cited sources change a lot even when you ask the same question again on the same day. Under the Don't Measure Once paper's experimental conditions, cited-source Jaccard between same-day runs was 0.32–0.43, and the standard error of a per-brand detection rate fell from 0.246 at two runs to 0.081 at seven.
Which of the three definitions does TRAIL Search measure?
It records mentions, citations, and share of voice inside answers separately for each engine, and labels each result by run count as a preliminary estimate (fewer than 3 runs), medium confidence (3 to 6), or high confidence (7 or more). Cells it could not measure stay unmeasured rather than zero. A source's actual contribution is not part of what it currently measures.
References
- [1]Aggarwal et al., "GEO: Generative Engine Optimization", KDD 2024
- [2]Nematov et al., "Source Attribution in Retrieval-Augmented Generation", arXiv:2507.04480 (2025)
- [3]Schulte, Bleeker & Kaufmann, "Don't Measure Once: Measuring Visibility in AI Search (GEO)", arXiv:2604.07585 (2026)
- [4]Puerto et al., "C-SEO Bench: Does Conversational SEO Work?", NeurIPS 2025 Datasets & Benchmarks
- [5]GEO project, code, and GEO-Bench (GEO-optim/GEO)
- [6]C-SEO Bench dataset (Hugging Face)
Summary
- Saying a brand was 'visible' in an AI answer translates into at least three values: visibility (how much and how early), contribution (did it actually shape the answer), and stability (does it hold up when repeated).
- The GEO paper's visibility metric measures the word count of citing sentences relative to the whole response, and decays position exponentially after normalizing by the number of sentences in the response.
- Shapley source attribution averages how much adding a document raises the log-likelihood of the original answer. The paper's experiments showed position bias and undervaluation of documents needed for synthesis.
- Under the Don't Measure Once paper's experimental conditions, cited-source Jaccard between same-day repeats was 0.32–0.43, similar to across days. Most of the wobble is model randomness.
- High visibility can coexist with low contribution, and both wobble when measured once. A report that says 'GEO is working' should first say which definition it measured.
Check this topic against your own brand
TRAIL Search measures how ChatGPT and Perplexity answer your customers' questions and finds and fixes where the brand is missing, turning this post into a diagnosis. Start with 10 questions, no card required.
More posts

The prompt you track is not the query AI actually searched
AI visibility reports score each tracked prompt, but retrieval and citation happen on fan-out sub-queries. What year and English injection mean for measurement.

Which content types AI cites, by search intent
The content type AI search cites flips with intent: articles lead informational prompts (45.5%), listicles lead commercial ones (40.9%).

How are retrieval and citation different in AI search?
Retrieval and citation are different stages in AI search. Does ChatGPT cite YouTube? Why our content is missing from AI answers, and what zero citations mean.
Keep reading on this topic
- The LLMO trap: training data vs. AI answer citationsSeeding your brand in LLM training data and getting cited in AI answers are different mechanisms. How they differ, what you can measure, and why advice splits.
- What marketers worry about most in AI search: reporting dataMarketers worry less about vanishing from AI search than about lacking reliable reporting. We checked whether that differs by company size.
- Should AI visibility be measured through the API or the UI?The same prompts through the ChatGPT API and UI shift brand visibility 41% on average and change cited sources. What to check when choosing a measurement tool.
This post is part of the Analysis category, which collects all 34 posts on the topic. See all posts in Analysis