
AI Answer Brand Mentions Track Exposure More Than Page Fit
A stage model of brand visibility estimated from 34,960 unbranded GPT and Gemini observations
When do AI answers mention a brand? In 34,960 GPT and Gemini runs, answers with neither its own domain nor a branded query named it 2.8–3.8% of the time.
Whether an AI search answer mentions a brand moved far more closely with whether the engine pulled in that brand's own domain as evidence than with BM25-measured page fit. Across 34,960 observations of unbranded prompts run repeatedly through GPT and Gemini, answers with neither the brand's own domain nor a branded search query mentioned the brand 2.8% of the time on GPT and 3.8% on Gemini [1]. When the own domain was cited, the rates were 49.0% and 58.4% [1].
We read the paper on arXiv (v1), including Tables 1–4, and checked every number against the table values. All figures come from the paper's data, and we did not reproduce them [1]. The author founded a company that builds AI search measurement software, and the paper discloses that all of the analyzed data is that company's tracking data [1]. The raw data is not public; only aggregate results and figure code are. We also build AI search diagnosis tools.
What makes an AI search answer mention a brand?
The study splits the conditions into four stages [1]. Does a page fit the real request (Match), does the engine pull that evidence in (Exposure), does it pick the brand from that evidence (Selection), and is there a prior path where the brand appears without live evidence (Prior)? The author sums it up with a mnemonic: "visibility ≈ match × exposure × selection + prior" [1].
The multiplication is deliberate. On the live retrieval path, one weak early stage blocks the result no matter how strong the later ones are. If the engine never retrieves a page, it never gets a chance to be selected [1]. The prior is added to guard against the opposite mistake. If every recommendation is assumed to come from live search, established brands that keep getting mentioned without citations go unexplained.
The contrast with traditional search is clear. The author frames traditional search as "relevance + authority" and generative search as "match + exposure + selection + prior" [1]. Ranking results is not the end; there is an extra answer-composition step that decides which retrieved pages make it into the answer.

Figure 1. Brand mention rate by retrieval evidence state.
How did the study measure brand mentions?
The study combines four datasets that each look at a different stage, not a single experiment [1]. The core is a large panel collected from 75 tracking projects between June and September 2026: 34,960 observations of 2,854 distinct prompts run repeatedly through GPT and Gemini [1]. Prompts that already contain the brand name were removed, since a mention there is expected.
Each observation recorded two live signals [1]. Whether the engine's stored source URLs included the target brand's own domain (own-domain exposure), and whether the engine generated a search query containing the brand name from an unbranded prompt (branded fan-out). Fan-out is the engine expanding one question into several sub-queries. Google's documentation says AI Overviews and AI Mode may use a "query fan-out" technique, "issuing multiple related searches across subtopics and data sources" [2].
The other three datasets look at earlier stages separately. Eighty real user prompts show when fan-out happens, 199 prompts and 275 site pages from one organization show how page fit relates to exposure, and 160 runs of 16 prompts repeated ten times for another organization show mentions without citations [1]. Page fit was scored with BM25, a standard search relevance score.
| Data | Size | Stage observed |
|---|---|---|
| Real-prompt fan-out | 80 unique prompts | Search activation |
| Organization A | 199 prompts, 275 site pages | Match → exposure → selection |
| Organization B | 16 prompts × 10 runs = 160 | Prior path |
| Large panel | 34,960 observations, 75 projects | Full model validation |
Table 1. The four datasets combined in the study [1].
Do ChatGPT and Gemini mention a brand more when its own domain is cited?
In the observations, yes. Answers that pulled in the own domain mentioned the brand far more often, and with a branded fan-out as well the rates were 91.4% on GPT and 100% on Gemini [1]. Compared with no signal at all, own-domain exposure alone corresponded to a 17.4x mention probability on GPT and 15.4x on Gemini [1].
| Retrieval evidence state | GPT observations | GPT mention rate | Gemini observations | Gemini mention rate |
|---|---|---|---|---|
| Neither | 15,524 | 2.8% | 13,801 | 3.8% |
| Own domain cited only | 1,769 | 49.0% | 3,415 | 58.4% |
| Branded fan-out only | 59 | 64.4% | 33 | 84.8% |
| Both | 128 | 91.4% | 231 | 100.0% |
Table 2. Brand mention rate by retrieval evidence state, unbranded prompts only [1].
This could be an illusion created by how well known each brand already is. The study checked by looking inside cells where the same organization and prompt were measured repeatedly [1]. With organization and prompt held fixed, the common odds ratio for own-domain exposure was still 15.3 on GPT and 29.7 on Gemini [1]. Differences that do not change over time, such as brand identity or prompt wording, do not explain it.
The two engines did not always pick the same brands either. They agreed on whether to mention the brand 84.7% of the time on the same prompt, with 1,999 Gemini-only mentions and 680 GPT-only mentions [1]. One engine's result cannot stand in for the rest.
Does a well-optimized page get cited in AI search?
A page that fits the question is a starting condition, but it did not decide citation on its own. Page fit predicted own-domain citation with AUC 0.641 on Gemini and 0.545 on GPT, close to chance [1]. On Gemini, prompts in the top quarter of page fit were cited 50.0% of the time versus 22.0% for the bottom quarter [1].
The exposure effect is much larger. On the 30 prompts where GPT cited the own domain, the brand was mentioned every time; on the other 169, only 9.5% [1]. Gemini was at 73.8% versus 8.2% [1]. Page fit roughly doubled the citation rate, while exposure split the mention rate by about nine to eleven times.

Figure 2. The exposure effect among well-matched prompts.
The best-matched prompts tell the same story. Of the 50 top-quarter prompts on Gemini, the 25 with own-domain exposure had an 84% mention rate and the 25 without had 8% [1]. On GPT, all 8 exposed prompts mentioned the brand, and the 42 unexposed ones were at 11.9% [1]. The author's advice is to stop leaning on a single "AI optimization score" unless it states which stage it measures.
How do brands get mentioned without being cited?
Some brands keep getting mentioned even when their own site is not cited. Across 160 runs of 16 prompts on ChatGPT, Competitor 1 was mentioned 74 times while its own URLs were cited 7 times [1]. Competitor 2 had 65 mentions and 5 citations [1].
Answers with no citations at all make it clearer. On 4 prompts that produced no citations in all ten runs, 40 runs in total, Competitor 1 was mentioned 30 times, Competitor 2 28 times, and the measured organization never [1]. On a prompt where that organization's dedicated page ranked in Bing's top 30, ChatGPT still mentioned it 0 times out of 10 [1].
The author does not call this path pretraining memory. The data cannot separate third-party pages, unobserved retrieval, and model knowledge, so it is called only a "prior-compatible path" [1]. At least in this organization's 160 runs, own-domain citation counts could not explain brand mention counts.

Figure 3. The brand visibility stage model and its predictive power.
How do I get my brand mentioned in AI answers?
First find which stage is blocking. The author's diagnostic order is this [1]. If match is weak, build or fix a page that answers the real question. If match is strong but exposure is weak, writing more on the same site may not help; look at distribution instead, such as indexing, crawlability, third-party coverage and comparison pages that put the evidence in front of the engine's retrieval step. If exposure is fine but the brand drops out of the answer, study how the engine chooses among its evidence.
The stage split mattered for prediction too. Predicting the next run's brand mention from prior mention history alone gave AUC 0.937 on GPT, live retrieval signals alone 0.880, and both together 0.963 [1]. Gemini was at 0.917, 0.840 and 0.942 [1]. History and live evidence complement each other; neither replaces the other.
In practice, the unit of measurement comes first. Of 23 commercial prompts, 78.3% triggered fan-out, against 3.6% of 55 informational prompts [1]. As in the example where a mascara request fanned out into both "how to choose" and "best" queries, a brand competes against the evaluation space the engine builds, not just the user's original sentence. This is the same issue as the gap between tracked prompts and the queries engines actually run.
Limitations
The most important limit is that this is an observational study. Own-domain citations and brand mentions are produced together in the same answer process, so the data cannot tell whether citation leads to mention or mention leads to citation [1]. The author states plainly that adding a URL to a source list would not raise mentions tenfold.
The data source also needs care. The raw data is the author's company tracking data and is not public [1]. Organization A's site was crawled on September 19 while visibility was measured on July 6, so page fit is a retrospective approximation. Own-domain citation captures only part of evidence exposure, and third-party pages may matter more. The fan-out data covers just 80 prompts.
The outcome unit has limits too. A brand mention is not the same as a positive recommendation, rank, click or purchase [1]. Baselines also shift with engine and model version: the no-signal mention rate was 1.8% on GPT-5 nano and 6.3% on GPT-5.4 [1]. This fits a recent survey that treats GEO as a stochastic, partially observable pipeline rather than a single ranking task [3]. Read the numbers as direction, not constants.
At TRAIL
One of the values TRAIL Search measures every day is the same kind of signal as this study's exposure. For each engine we collect the URLs an answer cites, and repeatedly record whether the brand's own domain is among them and whether the brand is mentioned. Cited sources are split into owned, wiki, YouTube, reviews and communities, media and listicles.
We do not observe an engine's internal fan-out queries from outside. This study observed part of that through vendor instrumentation.
Reading it made one point clear: looking at own-domain exposure against brand mention as four cells separates a page-fit problem from a distribution problem and an answer-selection problem. Source mix continues in our posts on where AI gets brand information and the difference between retrieval and citation.
Frequently asked questions
What makes an AI search answer mention a brand?
In this study's observations, the biggest divider was whether the engine pulled the brand's own domain in as evidence. Across 34,960 unbranded prompt observations, answers with neither the own domain nor a branded search query mentioned the brand 2.8% of the time on GPT and 3.8% on Gemini, versus 49.0% and 58.4% when the own domain was cited. These are observations, not causal effects.
Does a well-optimized page get cited in AI search?
Page fit is a starting condition, not enough on its own. Even for prompts in the top quarter of page match, Gemini mentioned the brand only 8% of the time when its own domain was not exposed. How well page fit predicted citation also differed by engine: AUC 0.641 on Gemini and 0.545 on GPT, close to chance.
Do ChatGPT and Gemini recommend the same brands?
On the same prompt, the two engines agreed on whether to mention the brand 84.7% of the time. In the rest, Gemini alone mentioned it in 1,999 cases and GPT alone in 680. That is why each engine needs its own measurement.
How do I get my brand mentioned in AI answers?
Start by finding where it breaks. If no page fits the real question, fix the page. If a page fits but the engine never retrieves it, look at distribution such as indexing and third-party coverage. If it is retrieved but left out of the answer, study how the answer is composed. Mentions without citations are a separate, slower-building brand base.
References
- [1]Benjamin Tannenbaum, "From Prompt to Recommendation: A Fitted Stage Model of Brand Visibility in AI Search", arXiv 2026
- [2]Google Search Central, "AI features and your website" (updated 2025-12-10)
- [3]Olivier Martinez, "Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023–2026)", arXiv 2026
Summary
- This stage model splits brand appearance in AI answers into page match, exposure, selection, and a prior path that works without live evidence.
- In 34,960 unbranded observations, answers with neither the own domain nor a branded query mentioned the brand 2.8% (GPT) and 3.8% (Gemini) of the time, versus 49.0% and 58.4% with an own-domain citation.
- Within repeated runs of the same prompt, own-domain exposure and mentions still moved together strongly (stratified odds ratios 15.3 and 29.7).
- Even top-quarter page match produced an 8% Gemini mention rate without exposure, and page fit predicted citation differently by engine.
- The author runs an AI search measurement company and the raw data is private, so read the numbers as direction.
Check this topic against your own brand
TRAIL Search measures how ChatGPT and Perplexity answer your customers' questions and finds and fixes where the brand is missing, turning this post into a diagnosis. Start with 10 questions, no card required.
The feature closest to this post is Keyword and question discovery.
More posts

Does AI Search Cite AI-Written Content? About 16% of Sources
Does AI search cite AI-written content? About 16% of sources four AI answer engines cited were AI-generated. How engines differ and what Google's policy says.

Where Does AI Get Brand Information? 85.7% Is Third-Party
Where does AI get brand information? Of 167,551 citations for 128 European brands, 85.7% were third-party sites. The top source shifts by language.

Why Your Business Doesn't Show Up in AI Recommendations
Why your business doesn't show up in AI recommendations: 85.6% of 4,776 Bali venues were never recommended. Documentation set entry; star rating set rank.
Keep reading on this topic
- Adding a diversity reward to query fan-out: Google's R4TR4T trains query fan-out on groundedness, diversity, and alignment rewards, then distills it into a small diffusion model. Is that diversity right for search?
- Which of two pages do AI answers cite? The 4 gatekeepersWhat does content need to get cited over a competitor in generative search? A 252K-run study finds four gatekeepers, and ChatGPT and Gemini differ.
- A survey of 45 GEO papers: what holds up and what doesn'tA survey of 45 GEO papers: four research shifts in three years, nine stages of AI visibility, why content-only optimization cuts citations, and what reproduces.
This post is part of the Papers category, which collects all 17 posts on the topic. See all posts in Papers