
How are retrieval and citation different in AI search?
YouTube is the domain ChatGPT fetches most, yet it almost never shows up as a citation in answers
Retrieval and citation are different stages in AI search. Does ChatGPT cite YouTube? Why our content is missing from AI answers, and what zero citations mean.
The sources AI search fetches most are not the sources it cites most in answers, so "our content doesn't show up in AI answers" hides two different diagnoses: "the engine doesn't fetch it" and "the engine fetches it but doesn't use it." The clearest example is YouTube. In one observation, YouTube was the domain ChatGPT retrieved most, yet across thousands of chats no inline citation of it inside an answer was observed [1]. Splitting the probability of citation into two stages, retrieval and selection, shows why this gap appears and what you should fix.
First, the scope and method of our evidence. The source-level observations come from Malte Landwehr of Peec AI, who shared his own observations in LinkedIn posts [1]. He turned trending stories on the Google News home page and topic pages into questions, ran more than 8,000 of them automatically over a month, and asked again every 6 hours until each 24-hour news cycle ended. This is vendor data that has not been peer reviewed, so by our standard its evidence grade is medium or below, and we do not use it to justify any scoring weight. Ranks and frequencies were not published, so this post does not invent ratios either. We cross-checked the structure in which retrieval and citation are different stages against the original RAG paper and the What Gets Cited paper [2][3]. This reflects what we knew as of October 2026, and we also disclose an interest: Peec AI is a competing vendor in our market.
Retrieval and citation are different stages
When an AI search engine gets a question, it first fetches documents, then picks some of them and attaches them to the answer as sources. The prototype of this structure is retrieval-augmented generation (RAG), where a retriever fetches external documents and a generative model writes an answer based on them [2]. In real products, one more selection step sits in between.
The What Gets Cited paper frames this point directly: because only a small subset of retrieved sources are cited, citation selection itself becomes the visibility bottleneck [3]. Raising your rank and winning a citation slot are different tasks.
Before retrieval there is also a condition of access. OpenAI runs a crawler that surfaces websites in ChatGPT's search results (OAI-SearchBot) separately from a crawler for model training (GPTBot), and states that sites that block the search crawler will not be shown in ChatGPT search answers [4]. The details are in OpenAI's crawler documentation. If access is blocked, the probability of retrieval approaches zero and the later stages never begin.
Does ChatGPT cite YouTube? Most retrieved, with no inline citation observed
In an observation of news queries, ChatGPT retrieved YouTube more than any other domain, yet no inline citation in its answers was observed, and sources split into two groups [1]. One group is retrieved often but rarely cited. YouTube was the most retrieved domain, and Facebook and Instagram were also retrieved often but rarely cited. Wikipedia, IMDB, and Reddit followed the same pattern.
The other group turns retrieval into citation. Reuters and AP dominated every category, even technology and entertainment, where wire services are thought to be weak. Health agencies such as the CDC, FDA, NIH, and WHO had an unusually high rate of being cited once retrieved. The same observation reported that news prompts went through web search 100% of the time [1].

Figure 1. The citation probability decomposition, with sources from the observation split into two boxes: "often retrieved, rarely cited" (YouTube, Facebook, Instagram, Wikipedia, IMDB, Reddit) and "retrieval turns into citation" (Reuters, AP, CDC, FDA, NIH, WHO). Ranks and frequencies were not published, so they are not drawn to scale.
The observer explained the difference as source curation to keep publishers on side and avoid misinformation, plus a preference for wire services that are easy to license [1]. That is the observer's guess, not something the engine has confirmed. Whichever explanation is right, though, the YouTube case is far from a problem you can solve by rewriting video descriptions.
Why doesn't our content show up in AI answers: retrieval or selection?
Our content does not show up in AI answers because either the retrieval probability or the selection probability is low, and the probability of being cited can be written as the product of the two events. This is the standard definition of conditional probability, and it only assumes that citation happens among retrieved documents.
is the probability that the engine fetches this document to answer the question. is the probability that, among the fetched documents, it picks this one and attaches it as a source. The two terms are different jobs, and because they multiply, if either one is near zero, raising the other does not raise the total.
| Low term | Symptom | What to do first |
|---|---|---|
| The engine does not fetch the document as a candidate | Crawler access, search visibility, vocabulary that overlaps with the question. Polishing the body text has no effect | |
| The engine fetches it but does not pick it as a source | Title, structure, wording actually used in the question, paragraphs an answer can be lifted from directly | |
| Both high but not cited | The engine does not treat that source type as citable | Instead of editing the document, check which other source types actually get cited for that question |
Table 1. Diagnoses and first actions depending on which of the two terms in citation probability is low. A classification by TRAIL Labs.
We covered how to fix documents that get fetched but not picked in Which pages does ChatGPT cite?, from the angle of question-page fit and how easily an answer can be extracted. Which of two candidates already in the context gets picked is covered in Which of two pages will AI cite?, based on a 252,000-run controlled experiment.
An experiment that measured only the selection term
The second term of the decomposition, , can be measured on its own with a controlled experiment. The What Gets Cited researchers placed two documents directly into the answer context as the only search results and instructed the model to cite exact URLs from those results only [3]. Because the experimenters fixed retrieval, is 1 for both candidates, and all of the remaining variance belongs to the selection stage.
Under that setup they attached 18 content factors to six commercial LLMs and asked 252,000 times, and four factors were classified as gatekeepers: topic match, a stated price, a recent timestamp, and a front slot in the context [3]. Only after these four were cleared did secondary factors such as specs, comparison information, and supporting evidence decide the winner, with odds ratios in the 2.1–243 range, while formatting differences such as paragraph density showed no consistent effect. All figures are under the paper's experimental conditions.
These results read accurately as the list of fixes for the second row of Table 1, the "fetched but not picked" case. They do not answer the problem in the first row. The paper itself notes as a limitation that real products retrieve many more documents, and that retrieval rank and presentation order get mixed up with content effects [3]. What you can raise by fixing a document is the second term; the first term is bound more tightly to conditions outside the document.
The YouTube case sits closest to the third row
YouTube had a very high probability of retrieval, but no citation was observed. In terms of the decomposition, is near zero. If that term is low because of the quality of individual videos, you could raise it by fixing descriptions and structure. But if it is consistently low for the whole source type, it is closer to the engine not treating that source as citable.
This distinction matters in practice because it changes where you invest. If there is a set of queries where work on a YouTube channel feeds the engine's information gathering but never turns into source attribution in answers, then winning citation slots for those queries has to happen in other source types. How engines differ in picking and displaying sources is covered separately in How ChatGPT, Perplexity, and Gemini cite differently.
Which fetched documents actually influenced the answer is a question at yet another level. Documents that carry a source label and documents that contributed to the content of the answer are not always the same, and there is separate research on computing this attribution [5]. We covered that method in Calculating source attribution with Shapley values.
When the engine and query change, YouTube's place changes too
Reading this as "YouTube is never cited anywhere" is an overgeneralization. In a separate analysis of 30 million sources published by the same observer, limited to the US and combining ChatGPT, Google AI Mode, Gemini, Perplexity, and AI Overviews, the most cited domains ranked Reddit, YouTube, LinkedIn, Wikipedia, and Forbes [1]. Combine several engines and several industries, and YouTube sits near the top of citations.
The two observations do not contradict each other. ChatGPT fetching YouTube for news queries without using it, and YouTube being cited often on other engines and other query types, can both hold at once. The observer himself noted as limitations that the 30-million-source analysis aggregates all industries, so it does not indicate importance in any specific industry, and that he did not publish percentages or counts [1].
So when you look at source types, look at the engine, query type, and time of measurement together. "YouTube doesn't get cited" and "YouTube is the second most cited" can both be true, and neither may be the answer for your category. Only a citation list you collect yourself, by engine, with your category's questions, gives you that answer.
What does zero citations mean in an AI visibility report?
Zero citations in a visibility report do not mean zero retrieval. Retrieval is not something you can observe from the outside. Which documents an engine fetched internally shows up only as hints on some screens, and it is not published as a number external measurement tools can count consistently. So even if a report shows zero citations for your page, that number alone cannot tell you whether it means "not fetched" or "fetched but not used."
The What Gets Cited paper proposes a diagnostic flow in which, if a brand is absent from citations entirely, the bottleneck is retrieval and the action is to improve SEO [3]. It is a reasonable starting point, but the YouTube observation shows that this rule has exceptions. Even the most retrieved source can be completely absent from citations. Blame the retrieval stage based on absence alone, and you end up fixing the wrong thing.
At TRAIL, we record only citations. We log which URLs were attached as sources, by engine, by question, and by check time. Retrieval cannot be observed from the outside, so we make no claims about it. In turn, we do not read zero citations as zero retrieval. The principle that unmeasured vs. zero are not the same number applies here too.
The same measurement shifts with model versions and reruns
Citation behavior by source is not a constant. The same observer recorded several cases where citation behavior changed each time the model version moved up [1]. For example, he reported that pricing pages were cited more often after the GPT-5.6 update. When versions change, which sources fall on the "fetched and used" side can change too.
Even within the same version, the swings are large. The Don't Measure Once study reported that asking the same question again immediately on the same day yielded only 0.32–0.43 overlap in cited sources (Jaccard) [6]. That is under the paper's experimental conditions, and the original is on arXiv. Draw conclusions about source-type patterns from a single citation list and you will mistake that day's variance for a pattern. Measure source-level numbers several times and report how many runs produced them.
Limitations
The source observation in this post is a single observation on news topics, in English, centered on ChatGPT, and we have not reproduced it [1]. Other query types such as product comparisons or local search produce different source mixes. In other analyses by the same observer, the types of content that get cited also varied widely by industry and intent.
Because ranks and frequencies were not published, we cannot quantify how large the gap between "most retrieved" and "no citation observed" is. The explanations for why YouTube is not cited (source curation, licensing preference) are the observer's guesses. The 30-million-source analysis is also limited to the US and does not cover Korean-language queries or Korean search surfaces. The What Gets Cited controlled experiment pitted two B2C review blogs with anonymized brands against each other, so it cannot be compared directly with how an engine chooses between wire services and video platforms for news queries [3]. The decomposition is the definition of conditional probability, so on its own it does not tell you the size of either term. Measuring each term fully is possible only from the side that can observe the retrieval stage, that is, inside the engine.
A practical order of checks
If an AI visibility report you are looking at says "not cited," start by checking whether you can tell which case it is. Here is the order. First, confirm that access for search crawlers is open. Second, collect, engine by engine, which source types actually get cited for that question. Third, if your document belongs to a type that gets cited but is missing, fix the document; if your source type itself is not cited for those queries, change the plan toward building a presence in the types that are. Confirm the numbers at each step with repeated measurement.
Frequently asked questions
How are retrieval and citation different in AI search?
Retrieval is the stage where the engine fetches documents to answer a question, and citation is the stage where it picks some of those documents and attaches them to the answer as sources. Not every fetched document gets cited, so a source being retrieved often is no guarantee that it gets cited often.
Doesn't ChatGPT cite YouTube?
In one observation where a vendor researcher ran more than 8,000 automated queries on news topics over a month, YouTube was the most retrieved domain, yet no inline citation of it was observed across thousands of chats. It is a single observation on news, in English, centered on ChatGPT, so results can differ for other query types and model versions.
If our content doesn't show up in AI answers, what should we fix first?
First figure out which of the two diagnoses applies. If the engine never fetches the document, fixing the document will not help, and you should look at search visibility and crawler access first. If the engine fetches it but does not pick it, fixing the title, the structure, and the wording actually used in the question is the right move.
Do zero citations mean zero retrieval?
No. Retrieval is not something you can observe directly from the outside, so zero citations cannot be the same number as zero retrieval. Read unmeasured vs. zero as two different things.
References
- [1]Malte Landwehr (Peec AI), LinkedIn posts: observation of 8,000+ automated ChatGPT chats based on Google News trending stories (2026, not peer reviewed)
- [2]Lewis et al., "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks", NeurIPS 2020
- [3]Vishwakarma, Kumar & Jamidar, "What Gets Cited: Competitive GEO in AI Answer Engines", SIGIR 2026
- [4]OpenAI, "Overview of OpenAI Crawlers"
- [5]Nematov et al., "Source Attribution in Retrieval-Augmented Generation", arXiv 2025
- [6]Schulte, Bleeker & Kaufmann, "Don't Measure Once: Measuring Visibility in AI Search (GEO)", arXiv 2026
Summary
- The probability that a page is cited in an AI answer is the product of the probability the engine fetches it and the probability it picks it from what was fetched. The two terms are different stages, and which one is low flips what you should do.
- In an observation of 8,000+ news queries, YouTube was the most retrieved domain but no inline citation was observed, while Reuters and AP dominated every category. It is a social media observation, so its evidence grade is medium or below.
- The single sentence 'we are not showing up in AI answers' hides two diagnoses: 'it is not fetched' and 'it is fetched but not used.' A single visibility score does not separate them.
- Retrieval cannot be observed from the outside, so do not read zero citations as zero retrieval. Unmeasured and zero are different numbers.
- Citation behavior by source changes with query type and model version, so read the numbers as direction rather than constants, and confirm them with repeated measurement.
Check this topic against your own brand
TRAIL Search measures how ChatGPT and Perplexity answer your customers' questions and finds and fixes where the brand is missing, turning this post into a diagnosis. Start with 10 questions, no card required.
More posts

How is AI answer visibility defined and measured?
How is AI answer visibility defined and measured? Three metrics with formulas and paper figures: visibility (GEO), contribution (Shapley), and stability.

The prompt you track is not the query AI actually searched
AI visibility reports score each tracked prompt, but retrieval and citation happen on fan-out sub-queries. What year and English injection mean for measurement.

Which content types AI cites, by search intent
The content type AI search cites flips with intent: articles lead informational prompts (45.5%), listicles lead commercial ones (40.9%).
Keep reading on this topic
- The LLMO trap: training data vs. AI answer citationsSeeding your brand in LLM training data and getting cited in AI answers are different mechanisms. How they differ, what you can measure, and why advice splits.
- What marketers worry about most in AI search: reporting dataMarketers worry less about vanishing from AI search than about lacking reliable reporting. We checked whether that differs by company size.
- Should AI visibility be measured through the API or the UI?The same prompts through the ChatGPT API and UI shift brand visibility 41% on average and change cited sources. What to check when choosing a measurement tool.
This post is part of the Analysis category, which collects all 34 posts on the topic. See all posts in Analysis