
The LLMO trap: training data vs. AI answer citations
One checklist bundles two different mechanisms into one
Seeding your brand in LLM training data and getting cited in AI answers are different mechanisms. How they differ, what you can measure, and why advice splits.
Seeding your brand in LLM training data and being cited as a source in AI answers are different mechanisms, and an AI search checklist that bundles them together leads you to pick the wrong fixes and the wrong way to verify them. The training path enters the model's weights at training time and stays fixed until the next run, and you cannot confirm from the outside that it moved. The retrieval path fetches documents the moment a question comes in, writes an answer, and attaches sources, so it operates now, changes per query, and lets you count the cited URLs. Almost everything GEO research measures today is the second one.
First, the scope and method of our evidence. Our starting point is the explanation in an LLMO guide that was widely read in early 2024. Some of that guide's technical statements are not accurate, so we did not quote them, and we did not list manipulation techniques. We cross-checked the distinction between the two mechanisms against the definitions of parametric and non-parametric memory in the original RAG paper and OpenAI's crawler documentation [1][2]. The deleted-article citation case is primary data from a vendor researcher's LinkedIn posts with no peer review, so by our standard its evidence grade is medium or below [6]. This reflects what we knew as of October 2026, and we also disclose an interest: we run an AI search measurement product.
AI search optimization checklists bundle two paths into one
The items on AI search optimization checklists are familiar. Get your brand onto Wikipedia and Reddit, list it on database-style sites, place an article with a large publisher or send out press releases. Most of it is reasonable work.
What does not hold up is the part that explains why the list works. A guide widely read in early 2024 explains it like this: if you could inject a million articles that mention a brand in the context of running shoes into the training data, you could make the model learn that association. This explanation bundles two different things into one. One is the training corpus; the other is retrieval and citation.
Names like LLMO, GEO, and AIO are defined a little differently across the industry, so which name you use does not solve this problem. Instead of the name, look at which mechanism each item assumes it works through. The same "get onto Wikipedia" becomes an instruction you cannot verify if you read it as "get into the training data," and an instruction you can verify by counting if you read it as "get into the documents answers cite today."

Figure 1. A TRAIL Labs diagram comparing Path A (training corpus) and Path B (retrieval and citation) on four rows (when it operates, how it varies, the scale required, and whether it is observable), along with the recommendation each one leads to ("publish everywhere" and "fix the documents answers pull from").
The training path: associations enter the model's weights
On the training corpus path, the association between a brand and a topic enters the model's weights at training time. The original RAG paper calls knowledge stored in model parameters this way parametric memory, and distinguishes it from non-parametric memory, knowledge drawn from an index of external documents [1].
This path has three properties. First, the effect stays fixed until the next training run. It is not a path where something you write today changes tomorrow's answer. Second, the scale a single company would need to meaningfully move a web-scale corpus is unrealistic. Third, there is no way to confirm from the outside that it moved. Training data composition and weights are not published, and you cannot look at a single answer and count which training document its content came from.
The retrieval path: documents are fetched and sources attached at query time
The retrieval and citation path operates the moment a question comes in. The engine fetches documents, writes an answer grounded in them, and picks some of them to attach as sources. RAG's non-parametric memory is the prototype of this path [1].
Its properties are the opposite of the training path. It operates now, it changes per query, and it is measured. Which URLs were attached as sources stays visible on the answer screen. The scale required is a single document at a time. Fix a document that answers actually pull from, and starting with the next query that change gets a chance to show up in results.
| Property | Training corpus path | Retrieval and citation path |
|---|---|---|
| When it operates | At model training time | At query time |
| Variance | Fixed until the next training run | Changes per query and per rerun |
| Scale required | Hard for one company to move | One document at a time |
| External observation | Not possible, nothing to count | Possible, count the cited URLs |
| Verifying the effect | No way to check | Check with citation logs by engine and question |
Table 1. A comparison of the two paths' properties. A table by TRAIL Labs that extends the classification in Figure 1 with a row on verifying the effect.
Which fetched documents get attached as sources, and how retrieval and citation diverge, is covered separately in How are retrieval and citation different in AI search?.
AI companies' documentation also treats training and retrieval separately
This distinction is not a classification we made up. It is already in AI companies' official documentation. OpenAI runs OAI-SearchBot, which surfaces websites in ChatGPT's search results, separately from GPTBot, which collects content that may be used to train its generative AI foundation models [2]. It also states that each setting is independent, so you can allow OAI-SearchBot to appear in search results while blocking GPTBot.
Where the line is drawn differs by company. Google's Google-Extended setting covers not only training Gemini models but also grounding, which passes content from the Google Search index to the model at query time, and the documentation states that it does not affect inclusion or ranking in Google Search [3]. The original is in Google's common crawlers documentation. You should not infer one company's boundaries from another's documentation, but treating the two mechanisms as separate is common to both.
What GEO research measures is almost entirely the retrieval path
The quantities GEO research measures today come almost entirely from the retrieval path: citation share, source distribution, and variance when the same question is asked again. The original paper that named GEO also measured how visible a source becomes inside generated answers when the retrieved source documents are edited [4].
Measuring variance is also a property of the retrieval path. The Don't Measure Once study reported that even when the same question was asked again immediately on the same day, cited sources overlapped (Jaccard) only 0.32–0.43 [5]. That is under the paper's experimental conditions, and the original is on arXiv. For a path that changes per query, it is natural that results swing from run to run. The training path is not an externally observable quantity to begin with.
The advice splits into "publish everywhere" and "fix the documents answers pull from"
The reason to separate the two paths is that the advice splits. Assume the training path and the instruction becomes "publish everywhere," because volume is the lever. Look at the retrieval path and the instruction becomes "find the documents answers actually pull from and fix those." The same checklist translates into two completely different quarterly plans.
What is interesting is that the targets the guide picked plausibly do work. The reason is different, though. If Wikipedia and large publishers work, it is because those documents are citable today. The explanation that they work by moving the training corpus has no way to be checked.
Once the distinction is clear, verification becomes possible too. For the retrieval path, you just count whether the page actually gets attached as a source, by engine and by question. For the training path, there is nothing to count. Attach "what would we count to know this worked?" to each checklist item, and it becomes obvious which path the item assumes. We laid out concrete ways to get fetched documents picked in Which pages does ChatGPT cite?.
Citations happen even when training crawlers are blocked
That the two paths are separate shows up most practically in crawler settings. Per OpenAI's documentation, a site owner can block the training crawler GPTBot and allow only the search crawler OAI-SearchBot, declining to provide training data while continuing to appear in ChatGPT search answers [2]. Conversely, block the search crawler and you drop out of ChatGPT search answers regardless of your training settings.
The observational side shows the same picture. In a vendor researcher's analysis of 30 million sources across several AI search engines, limited to the US, Facebook was the eighth most cited domain, and the observer himself noted that Facebook blocks most training crawlers [6]. A site that has closed off the training data path sits near the top of answer citations. The author also stated the limitations directly: the analysis aggregates all industries, and percentages and counts were not published.
So if you set crawler policy on the premise that "to be visible to AI, you have to get into the training data," the order gets tangled. If you want to be cited in answers, check search crawler access first, and decide separately whether to allow training crawlers as a matter of company policy on content use. That is the decision structure that fits the two mechanisms.
The autocomplete example shows a separate layer
One example the same guide uses is worth unpacking. It observes that autocomplete follows one company name with "whistleblower" and another with "whistle." That is a trace of those words co-occurring often in the training data. It shows, at most, that co-occurrence stays in the model.
Which source gets cited inside an answer is a different layer. Show the first and claim the second, and you skip a step. Between an association staying in the weights and your page getting attached as a source because of that association, there is a retrieval step.
There are also points where the two paths blend
The boundary is not always clean. One vendor researcher observed ChatGPT regularly citing a deleted Wikipedia article four months after it was deleted [6]. For the same article, Google, Bing, DuckDuckGo, and Brave returned no results, and Perplexity, Gemini, and Copilot did not cite it either. The leading explanation offered in the comments was parametric memory left over from training rather than live search, but that is a guess.
If this case holds, memory left over from the training path can reappear in an answer in the form of a source citation. The conclusion stays the same, though. Either way, what we can count is the sources attached to answers. Attributing the cause to training is a possible interpretation, but there is still no outside means to confirm it. The original is in Malte Landwehr's LinkedIn posts.
Read the checklist as a watch list instead of a to-do list
Setting the two apart also exposes the less comfortable half. The retrieval path is open to influence to the same degree it is measurable. The documents that become sources in answers are mostly not your site, and you do not control their standards. We covered research where a few lines on a page flipped LLM recommendations in Can AI recommendations be manipulated?.
So we read lists like this as a watch list rather than a to-do list. We look at which sources show up in answers for our category and whether that mix changes. Copying a technique and watching whether a technique is being used are different jobs.
At TRAIL, we state that what we measure is the retrieval path. We record which URLs were attached as sources, by engine, by question, and by check time, and we report how many runs produced each number. We make no claims about the training path. We see not pretending to measure what cannot be measured as the minimum condition for anyone who sells measurement.
Limitations
This post covers the distinction between the two mechanisms and does not measure the size of either path's effect. We did not quantify the scale needed to move a training corpus, and we did not test the effect of individual checklist items.
The guide we started from is from early 2024, so it does not reflect engine changes since then. The crawler documentation was checked in October 2026, and companies can revise it at any time. The deleted-article citation case is a single observation by a vendor researcher, the causal explanation is a guess, and we did not reproduce it [6]. The 30-million-source analysis behind the Facebook example is also limited to the US and aggregates all industries, so it does not transfer directly to Korean-language queries or specific industries. Putting crawler blocking and citation rank side by side is our interpretation; the observer did not analyze the relationship between the two. The variance figures in Don't Measure Once come from four Swiss German-language campaigns across four engines [5].
Next steps
Mark each item on the AI search checklist you are using with the path it assumes. For retrieval path items, measure several times whether the page gets attached as a source for your category's questions, engine by engine; for training path items, first ask whether the document itself offers any way to check. Before putting a quarter's budget on items with no way to verify them, it is safer to handle the countable items first. Even for retrieval path items, measure a baseline before touching anything, then measure again several times with the same questions after the change, so you can tell the change apart from that day's variance.
Frequently asked questions
If we put our brand into LLM training data a lot, will we be visible in AI search?
The training path stays fixed until the next training run, the scale a single company would need to meaningfully move the corpus is unrealistic, and there is no way to confirm from the outside that it moved. What gets attached as a source in AI search today is the retrieval path, where documents are fetched and cited at query time, and that path can be measured.
What is the difference between LLMO and GEO?
Definitions vary across the industry. The distinction that matters in this post is the mechanism, not the term. The training path, which plants associations in model weights, and the retrieval path, which fetches documents and attaches sources the moment a question comes in, differ in timing, variance, and measurability.
Does getting onto Wikipedia or large publishers work?
It plausibly does. The reason is different, though. Those documents are citable sources in answers today, and if so, you can count whether those pages actually get attached as sources.
How do you check whether a checklist item worked?
If the item assumes the retrieval path, you can measure several times whether the page gets attached as a source, by engine and by question. Items that assume the training path have nothing to count, so first ask whether the document itself offers any way to check the effect.
References
- [1]Lewis et al., "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks", NeurIPS 2020
- [2]OpenAI, "Overview of OpenAI Crawlers"
- [3]Google Search Central, "Google's common crawlers" (Google-Extended)
- [4]Aggarwal et al., "GEO: Generative Engine Optimization", KDD 2024
- [5]Schulte, Bleeker & Kaufmann, "Don't Measure Once: Measuring Visibility in AI Search (GEO)", arXiv 2026
- [6]Malte Landwehr (Peec AI), LinkedIn posts: observation of ChatGPT continuing to cite a deleted Wikipedia article (2026, not peer reviewed)
Summary
- The explanation commonly attached to AI search checklists bundles two different mechanisms into one: the training corpus path and the retrieval and citation path.
- The training path enters the weights at training time, stays fixed until the next run, and cannot be observed from the outside. The retrieval path operates per query, and you can count the cited URLs.
- AI companies' own documentation separates the two. OpenAI runs a search crawler and a training crawler separately and says their settings are independent.
- Assume the training path and the advice becomes 'publish everywhere'; look at the retrieval path and it becomes 'find the documents answers actually pull from and fix them.'
- The retrieval path is open to influence to the same degree it is measurable. That is why it is safer to read lists like this as a watch list rather than a to-do list.
Check this topic against your own brand
TRAIL Search measures how ChatGPT and Perplexity answer your customers' questions and finds and fixes where the brand is missing, turning this post into a diagnosis. Start with 10 questions, no card required.
More posts

How is AI answer visibility defined and measured?
How is AI answer visibility defined and measured? Three metrics with formulas and paper figures: visibility (GEO), contribution (Shapley), and stability.

The prompt you track is not the query AI actually searched
AI visibility reports score each tracked prompt, but retrieval and citation happen on fan-out sub-queries. What year and English injection mean for measurement.

Which content types AI cites, by search intent
The content type AI search cites flips with intent: articles lead informational prompts (45.5%), listicles lead commercial ones (40.9%).
Keep reading on this topic
- How are retrieval and citation different in AI search?Retrieval and citation are different stages in AI search. Does ChatGPT cite YouTube? Why our content is missing from AI answers, and what zero citations mean.
- What marketers worry about most in AI search: reporting dataMarketers worry less about vanishing from AI search than about lacking reliable reporting. We checked whether that differs by company size.
- Should AI visibility be measured through the API or the UI?The same prompts through the ChatGPT API and UI shift brand visibility 41% on average and change cited sources. What to check when choosing a measurement tool.
This post is part of the Analysis category, which collects all 34 posts on the topic. See all posts in Analysis