Is AI search visibility enough? Why we changed our question

Is AI search visibility enough? Why we changed our question

Showing up and getting through are different facts. What we ran into building a measurement engine, and what the papers showed first

AI search visibility is countable, but more of it doesn't mean your accurate information reached the answer. What measurement and three papers taught us.

By · TRAIL Labs Research
GEOAEOAI SearchMeasurementResearchTRAIL Labs

In AI search, visibility can be counted, but more visibility and your brand's accurate information actually making it into the answer are different facts. This post isn't a paper review. It's our own story. While building an engine to measure AI search visibility and reading the related papers, we kept coming back to the same question, and that question became the starting point for the research we're doing now.

First, the scope of our evidence. Every number here comes from public sources, the original papers, and each is marked as being from the paper's experiments [1][2][3][4]. We cross-checked links and numbers against the originals as of October 2026. On the product side, we only describe measurement rules already implemented in TRAIL Search, which we run ourselves. We're a company that runs an AI search measurement product, so we have an interest in the conclusion that "measurement needs to be more rigorous." Because peer review of our own research is in progress, we left its details and results out of this post.

Building AI search visibility measurement, we ran into three things first

Sending the same question to several AI engines, pulling brand mentions and citations from the answers, and computing share turned out to have more traps than we expected. Between getting a number and being able to trust it, we needed several rules.

First, rerunning the same question on the same day can change the cited sources a lot. The Don't Measure Once study, which measured this variation directly, observed 4 AI search engines across 4 campaigns for 45–46 days and reported that the overlap (Jaccard) between source sets from up to 10 consecutive same-day reruns was 0.32–0.43 (in the paper's experiments) [1]. Because that's almost the same range as the day-to-day overlap (0.34–0.42), the authors read most of the instability as coming from the randomness of model generation itself rather than from changes over time. In the same paper, the standard error of estimated brand appearance rates was 0.246 at 2 runs and only fell below 0.10 at 7 runs. A number measured once can't settle anything.

Second, putting a brand name in the question inflates share. Ask "how is Brand A?" and of course the answer mentions A. Mix questions like that in and you get a share that doesn't match the real competitive picture. So those questions had to come out of the share denominator.

Third, failing to measure and getting zero are different. If a cell where no engine response came back and a cell where the brand never appeared are both recorded as zero, the report tells a story that isn't true.

Problem we hitWhat happens if you leave itRule we built into TRAIL Search
Rerun varianceA single measurement gets mistaken for a trendConfidence badges by run count (under 3 runs preliminary, 3–6 medium confidence, 7 or more high confidence)
Brand-name questionsShare gets inflatedQuestions containing the brand name are excluded from the share denominator
Unmeasured vs. zero confusionUnmeasured cells look like no resultsUnmeasured and 0% are shown with different labels

Table 1. Problems we ran into building the measurement engine and the responses implemented in TRAIL Search. The confidence badge thresholds draw on the run-count analysis in Don't Measure Once [1].

As we set these rules one by one, numbers came out every day. But the more we built, the more one question remained: when citation rank goes up, does the answer the customer receives actually get better?

The GEO advice we heard in the field is right

We went to several seminars on GEO and AEO. The message was broadly similar. Build E-E-A-T and authority in your content, write on platforms AI often references, get your site's technical setup right, keep at it, and you'll show up. Check results daily or weekly, and track prompts designed around the customer journey.

All of that is right. These are things that genuinely need doing, and we also use TRAIL Search to design customer-journey questions and track them regularly. But one thing kept nagging at us as we listened. How do you confirm what these efforts actually changed? There was relatively little explanation of which change produced visibility, and how that visibility changed the answer customers receive.

We know better than anyone that measurement is hard. AI answers change every time, even for the same question. That's exactly why we think you need to be able to say what you changed, what you measured, and what you observed, all together. Whether results are good or fall short, you need to be able to examine why before you can make the next decision.

The papers show that the visibility race splits a fixed set of slots

Three papers from our reading series sharpened that question. Each, from a different angle, shows that a race to increase visibility doesn't automatically translate into better answers for users.

StudyExperimental setupResult relevant to this post (in the paper's experiments)
Original GEO paper (KDD 2024)Measured source visibility inside generative engine answers, by optimization techniqueWhen all sources were optimized at once, the "Cite Sources" technique left the top-ranked source at −30.3% visibility on average and the fifth-ranked source at +115.1%
C-SEO Bench (NeurIPS 2025 D&B)Measured gains while varying the share of participants using the same technique from 0% to 100%A zero-sum structure: gain per adopter shrinks as adopters increase and converges to zero at full adoption
RAG source attribution (2025)Allocated how much an answer relied on each document, using Shapley values and similar methodsAmong duplicate documents with the same content, the one placed first got higher attribution, and swapping the order flipped the bias

Table 2. Three studies that show the structure of the visibility race [2][3][4].

The original GEO paper ran a separate experiment where all sources are optimized at the same time [2]. With the "Cite Sources" technique, the fifth-ranked source's visibility rose 115.1%, while the top-ranked source fell 30.3% on average. The authors read this as opening opportunities for smaller creators, but the same table also shows that when someone goes up, someone else comes down. We broke the paper down in our review of the original GEO paper.

C-SEO Bench took this structure head on [3]. Measuring while increasing the share of participants using the same conversational search optimization technique, the average gain per adopter kept falling. The authors call it "a congested and zero-sum game" and write that even early adopters' gains converge to zero at full adoption. The same study also reported that placing a document in the first slot of the LLM context had a much larger effect than rewriting body text. In the Retail domain, first-slot placement improved citation rank by 2.77 positions on average, while the best rewriting technique managed 0.36 (in the paper's experiments). The benchmark code and data are public on GitHub and Hugging Face. There's more in our C-SEO Bench review.

The RAG source attribution study showed that "being cited" and "that document actually producing the answer" are different [4]. Allocating how much an answer relied on each document, it found that among duplicate documents with the same content, the one placed first consistently scored higher, and swapping the order flipped the bias too. The authors conclude that a single attribution score isn't enough and that you have to interpret the relationships between documents. We covered it in our RAG source attribution review.

Put the three together and one question remains. If everyone optimizes toward the same signals, whoever goes up pushes someone else down. At the end of that race, what's different about the answer the user receives? The brand spent money, but did the user get a better answer?

So we widened the measurement question to "did it get through?"

Alongside "how do we get cited more," we set a second question: "did the accurate information the brand holds actually reach the answer?" It's the difference between showing up and getting through. You can tell the first by counting mentions and citations. For the second, you have to check whether the content of the answer matches the brand's facts, and whether you can trace its evidence back to a source.

This idea is already in the product. TRAIL Search's AI drafts block numbers that have no source and are built to use only the evidence the brand has. Instead of promising visibility, we chose to produce drafts that contain verifiable information. The measurement integrity rules above follow the same principle. Only once you've separated what was observed from what wasn't can you ask what actually got through.

Our research is in peer review, so we're not putting numbers on its effect

We organized this direction into our own methodology and spent a long time setting hypotheses, designing experiments, and revising them. As a result, our submission to an international peer-reviewed conference in information retrieval is in progress. We want work we've validated internally to be evaluated through external review and shared more widely in the research community.

Until review is done, we won't put numbers on the effect. If someone who measures for a living talks up their own method's effect before it's validated, that's no different from the stories that never explained "what changed what." When results come in, we'll report them here the same way: which hypotheses we set, how we ran the experiments, and what was supported and what was rejected. We plan to publish rejected hypotheses as they are. We believe writing down what was wrong is part of validation.

What practitioners can check right now

There are things you can check today without waiting for research results. Every item below comes from published paper results or measurement rules already in operation.

  1. Don't judge from a single measurement. The Don't Measure Once authors recommend at least 7 runs per prompt per day for brand visibility monitoring, and a 2–4 week rolling window (in the paper's experiments) [1]. Mark measurements with few runs as preliminary.
  2. Remove brand-name questions from share. To see the competitive picture, build the denominator from questions a user who doesn't know your brand would ask.
  3. Keep unmeasured and zero in different cells. When the two get mixed in a report, the priority of improvement work goes wrong.
  4. Look at brand level and source level separately. In the same study, day-to-day overlap of brand mentions (Jaccard 0.45–0.59) was higher than source overlap (0.34–0.42) (in the paper's experiments) [1]. The authors conclude that brand-level aggregates are the more reliable campaign KPI, while tracking individual URLs is for diagnosing which content drives inclusion.
  5. Read the answer content along with whether you were mentioned. Beyond whether the answer mentioned your brand, check whether it got facts like price and specs right, and whether the cited page actually supports that sentence.
  6. Look at retrieval-stage ranking too. In C-SEO Bench, placing a document in the first context slot had a bigger effect than body rewriting techniques [3]. Before polishing body copy, check whether you're being retrieved well in the first place.

Limits

Here's what this post can't tell you.

  • It doesn't include the content or results of our research. Because review is in progress, we haven't disclosed hypotheses, methods, experimental design, or results. All numbers in this post come from other published studies.
  • Paper numbers come from each paper's experimental conditions. They were measured on specific engines, languages, and campaigns, so they don't transfer directly to Korean-language queries or other industries. They don't represent our customers' results either.
  • The seminar observations are our own experience. They aren't an industry-wide survey and don't refer to any specific event or company.
  • Product rules aren't the end of validation. Confidence badges and denominator rules make measurement less wrong, and tools that directly measure whether information got through to the answer are still at the research stage.

Wrap-up

In AI search, visibility can be counted, and even counting it takes discipline like repeated measurement and denominator rules. But published research shows the visibility race splits a fixed set of slots, and that being cited differs from contributing. So on top of "how visible were we," we started asking "did the brand's accurate information reach the answer," and that question is now going through peer review as research. We'd suggest checking whether your own AI search reports separate showing up from getting through.

Frequently asked questions

Is increasing AI search visibility the right goal for GEO?

Visibility is a starting point, but we don't think it's enough on its own. Visibility can be counted, but more visibility and your brand's accurate information actually making it into the answer are different facts. We changed our measurement question so the two are measured separately.

Why shouldn't you judge AI search visibility from a single measurement?

Because rerunning the same question on the same day can change the cited sources a lot. In the Don't Measure Once study, the overlap (Jaccard) between source sets from same-day reruns was 0.32–0.43, and the standard error of brand appearance rates only fell below 0.10 at 7 repeated runs (in the paper's experiments).

Does GEO stop working if everyone does it?

The research points that way. C-SEO Bench reported that as more participants use the same technique, the gain per adopter shrinks and converges to zero at full adoption. The original GEO paper also found that when all sources were optimized at once, the top-ranked source's visibility fell by 30.3% on average (in the paper's experiments).

When will TRAIL Labs publish its research results?

Our submission to an international peer-reviewed conference in information retrieval is in progress, so we won't put numbers on the effect before review is complete. When results are in, we'll report them on this blog the same way, including which hypotheses we set, what was supported, and what was rejected.

References

  1. [1]Schulte, Bleeker & Kaufmann, "Don't Measure Once: Measuring Visibility in AI Search (GEO)", arXiv 2026
  2. [2]Aggarwal et al., "GEO: Generative Engine Optimization", KDD 2024
  3. [3]Puerto, Gubri, Green, Oh & Yun, "C-SEO Bench: Does Conversational SEO Work?", NeurIPS 2025 Datasets & Benchmarks
  4. [4]Nematov et al., "Source Attribution in Retrieval-Augmented Generation", arXiv 2025

Summary

  • In AI search, visibility can be counted, but more visibility and your brand's accurate information actually reaching the answer are different facts.
  • Building a measurement engine, we first had to solve rerun variance, share inflation from brand-name questions, and confusion between unmeasured and zero.
  • The original GEO paper, C-SEO Bench, and RAG source attribution research show that the visibility race splits a fixed set of slots, and that being cited differs from contributing.
  • So we widened our measurement question from "how visible were we" to "did accurate information reach the answer," and we'll report results from this line of research after peer review.

Why TRAIL Labs works this way

How a researcher-founder grounds measurement and recommendations, and how the three products relate to the company, is laid out on the about page, and how we work day to day is written up separately.

More posts

Keep reading on this topic

This post is part of the TRAIL Labs category, which collects all 17 posts on the topic. See all posts in TRAIL Labs