0% vs unmeasured in an AI visibility report: how to tell

0% vs unmeasured in an AI visibility report: how to tell

The moment a report paints its blank cells as zero, 'absent' and 'unknown' become the same value

In an AI visibility report, 0% means measured and absent; unmeasured means no value. How to tell them apart, calculate share of voice, and set run counts.

By · TRAIL Labs Research
AI visibilityMeasurementShare of voiceGEOAEO

In an AI visibility report, zero and unmeasured are different numbers. Zero is an observation that you measured and the brand was not there; unmeasured is a gap where you never got a value for that cell. Most dashboards leave both cells blank or paint both as zero. That single collapse makes the number unusable for the decision it was meant to support. When someone reports "zero for this prompt group this quarter," the number alone cannot tell you whether you need to create content or fix the measurement path first.

Here is the scope of the evidence. The figures on repeated measurement come from the Don't Measure Once paper by researchers at the University of St. Gallen, which we checked against the original text, and all of them hold under the paper's experimental conditions [1]. The conditions under which engines produce no answer or no citations come from official Google and OpenAI documentation and public data from Pew Research Center [2][3][4]. We have a stake here: we run TRAIL Search, which measures AI search visibility. We use our own measurement rules as examples, but we do not claim those rules increase visibility or citations.

How to break down the causes when AI citations really are zero is covered in Zero AI search citations? Diagnose the cause step by step, and why you should split results by engine is covered in One AI visibility score is risky: ChatGPT isn't Perplexity. This post focuses on the step before both: deciding whether that zero was actually measured.

0% and unmeasured are different observations

Zero visibility is the result of a measurement that worked. The engine produced an answer, the answer was collected, and when we looked for a brand mention or citation, there was none. Unmeasured is a cell where measurement did not happen. Either there was no answer, or there was an answer without the information needed to judge it (such as a citation list), or collection itself failed.

The difference matters because the next action is the opposite. Zero is evidence that "this engine chose other sources instead of us for this prompt," so it is a starting point for content and source recommendations. Unmeasured means "we don't know anything yet," so the measurement conditions have to change first. If you add more content for an unmeasured cell and the cell is blank again next round, you have no way to check whether it worked.

Cell stateWhat actually happenedShown in reportNext action
Observed, mentionedThe answer was collected and the brand appearedValue (%)Maintain and expand per engine
Observed, not mentionedThe answer was collected and the brand was absent0%Review the competing cited sources and recommend fixes
UnmeasuredNo answer or citations, or collection failedUnmeasured (marked differently from a blank value)Check the measurement path and conditions first

Table 1. The three states a cell in an AI visibility report can have, and the next action for each.

Zero AI answer citations are often unmeasured cells

Unmeasured cells come more often from AI search working normally than from rare failures. Three paths are typical.

First, no AI answer is generated at all. Google Search Central states that AI Overviews are shown only when its systems determine they are additive to classic Search, and so they often don't trigger [2]. Usage data confirms this. When Pew Research Center analyzed 68,879 unique Google searches from the March 2025 browsing data of 900 U.S. adults, 12,593 of them, about 18%, produced an AI summary [3]. For a search with no AI summary, "our brand's AI Overview visibility" is undetermined, not 0%.

Second, the model skips web search, so the citation list is empty. OpenAI's web search tool documentation explains that, like any other tool, the model chooses whether to search the web based on the content of the input prompt. Under the default setting (tool_choice: "auto"), search is optional [4]. An answer produced without search has no cited URLs, so it cannot tell you whether your page was cited.

Third, collection fails or data is excluded for quality reasons. That is a problem on the measuring side, but in a report it shows up as the same blank as the first two cases.

A person pointing at a bar chart on a monitor; the chart shows a very short filled bar next to a hatched bar drawn only with a dashed outline

Figure 1. In the bar chart on the monitor, a very short filled bar sits next to a hatched bar drawn only with a dashed outline. The first stands for a cell with a small measured value, the second for a cell that could not be filled.

The repeated-measurement paper also handled zero and unmeasured separately

The Don't Measure Once authors applied this distinction directly in their analysis. The study ran eight prompts per vertical across four Swiss-German verticals (telecommunications, real estate sales, sporting goods, and consumer electronics) through four engines (ChatGPT, Gemini, Google AI Mode, and Perplexity) every day from January 24 to March 20, 2026 [1].

When comparing repeated runs from the same day, the authors kept only runs from which at least one citation was extracted. Leaving zero-citation responses in would inflate Jaccard values with false matches of the "both empty, so they agree" kind. 75.4% of runs passed this filter, and ChatGPT was lowest at 42.2%. The paper attributes this to ChatGPT's tendency to suppress web search on definitional queries [1]. In other words, more than half of ChatGPT's runs were unmeasured for citation-level judgments.

Daily collection had gaps too. Over the 45–46 day window, ChatGPT, Google AI Mode, and Perplexity had results for 38–44 days per vertical, while Gemini had only 22–26. January 30 was excluded from the analysis entirely because citation volume spiked to about twice the daily average [1]. Had those gaps been filled with zeros, Gemini would have been recorded as a far less visible engine than it was.

How to calculate share of voice in AI search: drop prompts that contain your brand name

Share of voice (SoV) in AI search is calculated over a prompt set with brand-name prompts removed. Separating zero from unmeasured is the same problem as setting the denominator honestly. The first rule we set before building any scoring was to remove prompts that ask about your own brand from the denominator. Asking an AI "What do you think of our brand?" and counting the brand when it shows up in the answer is not measurement. As a formula, it looks like this.

is the full prompt set, and is what remains after prompts containing the brand name are removed and flagged as contaminated. is the engine set, and per-engine values are shown before they are combined. is the set of brands that engine names when answering prompt .

The denominator is the number of prompt and engine pairs. If unmeasured cells slip in as zeros, the numerator stays the same while the denominator grows, and share of voice falls. The brand was not actually less visible, yet the number reads as a decline. Conversely, if brand-name prompts slip in, the numerator fills up almost automatically and share of voice inflates. Both errors come down to what you put in the denominator.

Unmeasured items leave the denominator instead of scoring zero

We apply the same principle to page audit scores. The second rule is to drop items we could not measure from the denominator instead of scoring them zero.

is the set of scoring items that were actually observed, is an item's weight, and is the share of that item's points the page earned. Our 100-point rubric breaks down as access 4, retrieval (lexical overlap) 13, lexical preservation 5, reranking 12 (chunk position 8 plus best chunk 4), citation rhetoric 6, intent alignment 4, completeness 4, structure 20, and authority 32. If a page has no target prompt, the retrieval and intent alignment items cannot be measured. If the page has no comparison context, the completeness item that looks at pricing and specs is not in scope either.

A worked example makes the difference clear. Take a hypothetical page with no target prompt that is not a comparison page. The items that drop out are retrieval 13, intent alignment 4, and completeness 4, so the maximum observable score is 79. If the page earns 60 of those 79 points, the formula above gives 60/79, or about 76 on a 100-point scale (our calculation). Score the unmeasured items as zero and the same page is recorded at 60. That 16-point gap comes from the scope of measurement and has nothing to do with page quality. Once a team starts working to close those 16 points, it spends time on things that never needed fixing.

How many times should you measure AI search visibility before trusting it?

Under the paper's experimental conditions, per-brand detection rates need seven or more repeated runs before the standard error drops below 0.10 [1]. A zero measured once is not yet a zero. The third rule is to show the run count first. Ask the same question again on the same day and the cited sources change a lot. In the Don't Measure Once paper, the Jaccard similarity of cited sources between runs repeated within 24 hours averaged 0.32–0.43 by vertical, the same range as the day-to-day figure of 0.34–0.42 [1]. Variation within a single day accounts for most of the observed instability. On a day-to-day basis, about 65% of cited sources turned over from one day to the next [1].

This variation shrinks as the sample grows. The standard error for repeated runs follows the familiar formula.

Runs (n)Standard error (SE)95% CI half-width (±)
10.3700.724
20.2460.483
30.1880.369
50.1230.241
70.0810.158
80.0620.121

Table 2. Standard error of the per-brand detection rate by number of runs (averaged across 1,216 per-brand series, under the paper's experimental conditions) [1].

The paper calls a single run essentially uninformative. A brand whose true detection rate is 50% could appear anywhere from −22% to +122% in a nominal 95% interval from one run. At seven runs the standard error falls to 0.081, below 0.10, and source-level coverage needed eight runs. For per-brand estimates, the authors recommend rolling aggregation over two to four weeks [1].

So before showing a value, we label each cell by run count: fewer than 3 runs is a preliminary estimate, 3 to 6 is medium confidence, and 7 or more is high confidence. A cell that showed zero once and a cell that showed zero in all seven of seven runs are not the same zero. The paper also notes that fewer runs suffice for brands that are either always or never cited [1]. A zero confirmed by repetition is solid evidence; a zero seen once is only a preliminary observation.

Evidence grade caps the weight

The fourth rule is about the honesty of the score itself. Each scoring item is tied to a published study and its evidence grade, and the grade caps how far that item can move the total.

is the evidence grade of item . An item backed only by evidence borrowed from another field (analogy) cannot exceed 5 of 100 points, and an item backed only by weak evidence cannot exceed 10. A scoring item with no study behind it cannot be created at all. Our chunk simulation, which cuts a page into 256-token chunks with 64 tokens of overlap, also uses a monotonically decaying position weight, strongest at the front. We tested a bonus for the conclusion section and did not ship it.

This rule grows from the same root as separating zero from unmeasured: do not put what you don't know into the score as if you knew it. Give a large weight to an item with weak evidence, and a zero on that item looks like a bigger problem than it is.

What to check first when a report arrives

When an AI visibility report arrives, check five things before reading the numbers.

  1. Check that zero and unmeasured are drawn differently. If they share the same color or the same blank, the report does not separate absent from unknown.
  2. Check whether brand prompts are in the denominator. If prompts containing the brand name went into the share-of-voice calculation, the number does not measure discoverability.
  3. Check how many runs produced each cell. Without a run count, you cannot tell whether a change from last month is real or model randomness.
  4. Check whether the prompt set changed. A rise during a period when the prompt mix changed is a change in composition, not an improvement. We attach a fingerprint to each prompt set to tell the two apart.
  5. Look at per-engine values before combining engines. If one engine's unmeasured cells blend into the overall average as zeros, they hide the other engines' results too. How to read results by engine is covered in detail in One AI visibility score is risky: ChatGPT isn't Perplexity.

Only a zero that passes all five is worth diagnosing. For the order in which to break that zero down into search visibility, crawling, and content structure, continue with Zero AI search citations? Diagnose the cause step by step. That measurement instability is a problem across GEO research is also clear from A survey of 45 GEO papers: what holds up and what doesn't.

Limitations

The figures in this post have clear boundaries. The Don't Measure Once results come from four Swiss-German verticals, eight prompts per vertical, four engines, and January to March 2026 [1]. The paper does not say whether Korean prompts or Korean search surfaces would show the same range of variation. The authors themselves leave open whether the same instability patterns hold in other languages and regional markets. The standard error thresholds also apply to per-brand detection rates and assume intermediate detection probabilities.

Pew Research Center's 18% figure is based on Google searches by U.S. users in March 2025 [3]. The share of searches that show an AI summary varies by query type and over time, and this post did not measure that share for Korean search.

Our four measurement rules are design choices for reading numbers honestly. They do not increase visibility, citations, or rankings. Those are outputs of models we do not control. What we report is what was observed, what could not be observed, and what moved since the last run.

Frequently asked questions

How should I read 0% versus an unmeasured cell in an AI visibility report?

0% is an observation: that engine and prompt pair was actually measured and your brand never appeared. Unmeasured means no value was obtained, because no answer was generated, no search or citation happened, or collection failed. 0% is a basis for content recommendations; unmeasured is a signal to fix the measurement path first.

How many times should AI search visibility be measured?

Under the Don't Measure Once paper's experimental conditions, the standard error of a per-brand detection rate was 0.370 at one run and 0.246 at two, and fell below 0.10 (0.081) at seven. Source-level coverage needed eight runs, and the authors recommend rolling aggregation over two to four weeks for per-brand estimates. That is why we label results by run count as preliminary, medium confidence, or high confidence.

Why shouldn't prompts that contain the brand name count toward share of voice?

If you ask 'What do you think of our brand?', the answer almost always mentions that brand. When such prompts sit in the denominator, share of voice rises regardless of how discoverable the brand actually is. So prompts that contain the brand name are removed from the denominator, flagged as contaminated, and shown separately.

What goes wrong if items we could not measure are scored as zero?

The reason you could not measure something starts to look like a weakness of the page. If a page with no target prompt gets zero on its retrieval and intent items, its score drops and the team fixes things that do not need fixing. Removing unmeasured items from the denominator distorts the judgment less.

References

  1. [1]Schulte, Bleeker & Kaufmann, "Don't Measure Once: Measuring Visibility in AI Search (GEO)", arXiv:2604.07585 (2026)
  2. [2]Google Search Central, "AI features and your website"
  3. [3]Pew Research Center, "Google users are less likely to click on links when an AI summary appears in the results" (2025)
  4. [4]OpenAI API docs, "Web search"

Summary

  • 0% is an observation that you measured and the brand was absent; unmeasured is a gap where no value was obtained. A report that draws both as the same zero cannot tell 'absent' from 'unknown'.
  • Unmeasured cells appear when no AI answer is generated, when the model skips web search so there are no citations, or when collection fails.
  • Remove brand-name prompts from the share-of-voice denominator, and drop unmeasured scoring items from the denominator instead of scoring them zero.
  • Under the Don't Measure Once paper's experimental conditions, same-day reruns produced cited-source Jaccard of 0.32–0.43, and the standard error of brand detection fell below 0.10 at seven runs.
  • When a report arrives, first check whether zero and unmeasured are drawn differently, how many runs produced each cell, and whether the prompt set changed.

Check this topic against your own brand

TRAIL Search measures how ChatGPT and Perplexity answer your customers' questions and finds and fixes where the brand is missing, turning this post into a diagnosis. Start with 10 questions, no card required.

More posts

Keep reading on this topic

This post is part of the Analysis category, which collects all 34 posts on the topic. See all posts in Analysis