
How ChatGPT understands YouTube: it reads the description
What a published retrieval snippet shows: the description becomes the video's body text, and the view count in the snippet isn't live
How does ChatGPT understand a YouTube video? It reads a text snippet with no transcript, and the description does most of the work. In the published example, the YouTube result ChatGPT received was a text snippet with the title, channel, publish date, views, likes and description, and no transcript.
ChatGPT doesn't watch YouTube videos, and it doesn't read their transcripts either. What it retrieves is a text snippet with the title, channel details, view count and the full description. So most of the text a model uses to decide what a video covers is the copy you typed into the description box under the upload form. This article starts from a published snippet example and adds two things that can be read from it: the numbers in the snippet are cached values frozen at retrieval time, and if you don't separate "retrieval" from "citation," you will misdiagnose a drop in YouTube citations.
The starting point is Peec AI's published analysis of the YouTube snippet, based on the reported ChatGPT retrieval leak[1]. Let us state the scope of the evidence first. The analysis rests on a reported leak and one real example. It is not a paper and has not been peer-reviewed. We have not reproduced the result. The view-count gap and the probability decomposition below are our own calculations and framing from the published example. We reopened and cross-checked the reports and official documentation cited here as of October 2026. We should also disclose an interest: Peec AI builds AI visibility measurement tools and works in the same market we do. We credit their finding as is and add one layer on top of it.
Does ChatGPT read YouTube transcripts? The snippet it receives has none
In the published example, what ChatGPT received when it retrieved a YouTube result was neither the video file nor the transcript, but a text snippet[1]. Here are the fields the snippet contained.
| Field | In the snippet? | Notes |
|---|---|---|
| Title | Yes | |
| Channel name | Yes | With follower count and verified status |
| Publish date | Yes | |
| Views and likes | Yes | Values at the time of retrieval |
| Full description | Yes | Including links, lists and timecodes |
| Transcript | No |
Table 1. The makeup of the snippet ChatGPT received for a YouTube result in the published example[1].

Figure 1. Lines of text continue below a video card on the screen, with an empty hatched dashed box beside them. Search, clock and link icons point toward the screen.
Whatever the video format, from the model's side a YouTube result is a text document: a few short lines of metadata with one long description attached. No matter how precisely something is explained in the video, if it isn't in the description, the model at this stage has no way to know it. Even for a video with a carefully uploaded transcript, the published example's snippet carried no transcript.
ChatGPT's search architecture makes this unsurprising. In an analysis of how ChatGPT built its own search index, Peec AI's Tomek Rudzki lists YouTube as a separate vertical index alongside general web, PDF, news, arXiv, Wikipedia, shopping and others[3]. According to the same article, Nick Turley, head of ChatGPT, said OpenAI began building its own search index in 2023, aiming to answer 80% of queries from it by the end of that year[3]. If YouTube results are retrieved through a separate path like this, the snippet above can be read as the format that path hands to the model.
Why the YouTube description becomes the body text in AI search
In the published example, the description was longer than all the other fields combined[1]. The description's share of the text a model holds about a video can be written like this.
Here is the length of the description and is the length of the snippet's full text. A description longer than all the other fields combined means this value is above 0.5. In other words, more than half the material the model uses to conclude what a video covers comes from one box: the description.
Seen this way, the description is not a box for extra information. The role that body text plays for a web page in search, the YouTube description plays in AI search. A description with only a few hashtags and a list of affiliate links is in roughly the same state as a web page with a title and no body. A description that states the video's subject in the first sentence and breaks what it covers into paragraphs and lists gives the model ample grounds to judge the video's content.
The view count in the snippet is a cached value from retrieval time
In the published example, the view count in the snippet was lower than the one shown on the YouTube page[1]. Both values are copied into Table 2 below, and our calculation of the gap as a share of the live value works out like this.
is the current value shown on the YouTube page and is the value in the snippet. By our calculation, the snippet value was a little more than one tenth (0.119) below the live value. The original analysis didn't address this mismatch, and the raw data for the calculation is the published example screen[1].
| Field | YouTube page | Snippet | Comparable? |
|---|---|---|---|
| Views | 64,410 | 56,729 | Yes (both exact) |
| Likes | 1.5K | – | No (rounded on the page) |
| Followers | 1.66M | – | No (rounded on the page) |
Table 2. YouTube page values versus snippet values for the same video. Only views were exact on both sides, so only views could be compared[1].
Likes and followers can't be compared the same way, because the YouTube page rounds them to figures like 1.5K and 1.66M. A gap this size on the one field that compares cleanly signals that the snippet holds a cached document saved at some point instead of a live read.
Why does that matter? If views and followers work as boosters at the retrieval stage, a video is scored on numbers from some point in the past rather than its current popularity. The longer the refresh cycle, the longer a video is scored at roughly its upload-day popularity. How long that lag runs can't be known from one example. What is clear is that the expectation "our views jumped recently, so AI search will reflect it soon" comes true only after that lag.
Why YouTube citations dropped in ChatGPT: retrieval and citation are different events
Being retrieved as a ChatGPT search result and being cited in the answer are different events for YouTube. The relationship between them can be written like this.
is the probability that a video is retrieved as a candidate at the search stage, and is the probability that, once retrieved, it is actually selected as a source for the answer. The citation share we see on a dashboard is the product of the two terms.
On September 25, 2026, Peec AI GEO researcher David Konitzny reported that YouTube's share of ChatGPT citations had fallen 91% compared with August. According to coverage by the French SEO outlet Abondance, facebook.com fell 88.1%, wikipedia.org 78.4% and forbes.com 71.7% over the same period, and these domains are still retrieved at the search stage but show up far less often among the final cited sources[2].
The 91% drop Peec AI observed is consistent with a collapse in either term[2]. If retrieval held and selection fell, the place to fix is the snippet. Working on the description, structure and terms gives the selection probability room to move. If retrieval itself fell, no amount of polishing the description will recover it. One number can't tell the two apart, and the two fixes point in opposite directions. We covered a separate observation, that YouTube ranks high in retrieval yet barely appears in citations, and how to diagnose it, in why the sources AI retrieves most differ from the ones it cites.
This distinction isn't limited to YouTube. For any page, "not visible in AI" splits into cases where it never made the candidate pool and cases where it made the pool but wasn't selected. If measurement can't separate the two, you end up applying one cause's fix to the other cause. We explained why a single combined metric is risky in why a single AI visibility score is risky.
To get cited in AI search, write the YouTube description like a landing page
Whichever term the drop came from, one fact holds: structure survives into the snippet. In the published example, timecodes and lists stayed intact in the snippet, and links were rewritten into reference tokens the model can point to rather than dropped[1]. That is why structuring the description like a web page matters on the selection side.
Chapter timecodes become statements in their own right about "what is covered when." According to YouTube Help on video chapters, to create chapters with timecodes in the description, the first timecode must start at 00:00, there must be at least three in ascending order, and each chapter must be at least 10 seconds long[4]. Follow those rules and chapters appear on the YouTube page, while in the snippet the chapter titles read as the video's table of contents.
Here is what to watch for when writing the description.
- State the subject in the first sentence. Say in one sentence what the video explains and for whom.
- Use the words people ask with. Instead of channel jargon or slang, use the phrasing viewers type into a search box or ask an AI.
- Write descriptive chapter titles. Rather than "Part 2," write something like "Pricing compared: monthly vs. annual plans" so each section says what you learn there.
- Break it into lists and paragraphs. Put anything that can be enumerated, such as topics covered, what you need and key figures, into lists.
- Don't fill it with only hashtags and affiliate links. Links survive as references, but a description made only of links tells the model nothing about the video.
Laying out how each part of the description carries into the snippet shows where the effort should go.
| Description element | In the snippet | How to write it |
|---|---|---|
| Descriptive sentences | Kept as is | Subject and audience in the first sentence |
| Lists | Kept as is | Topics covered, what you need, key figures |
| Timecodes | Kept as is | A descriptive chapter title for each section |
| Links | Rewritten as reference tokens | Only sources and related material |
| Transcript | Not included | Restate the key content in the description |
Table 3. How description elements carried into the snippet in the published example, and the writing principles that follow[1].
These principles aren't very different from the conditions under which web pages get cited. The need for the question and the page to match, and for the answer to be easy to lift out, is the same as what we covered in which posts ChatGPT cites. On YouTube, the only difference is that the "page" is the description.
Limits
The core observation in this article rests on a reported leak and one real example[1]. We have not reproduced it, and until the same result repeats across many videos and many queries, it should be treated as a direction instead of a fixed rule. The snippet format is ChatGPT's internal implementation and can change without notice.
The view-count gap in Table 2 is a comparison for one video at one point in time. The cache refresh cycle, the variation across videos, and how much views and followers actually weigh in retrieval ranking can't be known from this example. That is why we wrote "if they work as boosters" as a condition. The observation that the description is longer than the other fields also concerns this one example video, so we haven't confirmed how the share changes for videos with short descriptions. What does follow directly from Table 1 is that, for a video with an empty description, the only text the model has is the title and a few lines of metadata.
The 91% drop is an observation based on the data Peec AI tracks and has not been peer-reviewed[2]. There is no basis to assume its sample and question set match the markets we look at, especially Korean-language questions. This article also covers ChatGPT only. We didn't look at whether Perplexity or Google AI Overviews receive YouTube results in the same format.
The short version
Based on the published example, to ChatGPT a YouTube video is a text document made of a few lines of metadata and the description, with no transcript inside[1]. Since the description makes up more than half of that document, most of what the model knows about the video comes from the description. The view count in the snippet was a cached value, and a drop in YouTube citations can't be diagnosed until you separate whether it came from retrieval or selection. The surest thing you can do now is open the box the model reads in place of your video and look at what it says.
Frequently asked questions
Does ChatGPT read YouTube transcripts?
In the published snippet example, there was no transcript. What ChatGPT received was the title, the channel name with follower count and verified status, the publish date, views and likes, and the full description including links, lists and timecodes. This rests on a reported leak and one example, so it is too early to call it a general rule.
Why does the YouTube description matter for AI search?
In the published example, the description was longer than all the other fields combined. That means most of what a model concludes about a video's content comes from the description text. A description that states the subject, uses the terms people ask with, and is organized into chapters and lists is easier for the model to read.
What does a 91% drop in YouTube citations on ChatGPT mean?
It is a public observation by a Peec AI researcher that YouTube's share of ChatGPT citations fell 91% compared with August. A citation is the product of the probability of being retrieved and the probability of being selected once retrieved, so this one number can't tell you which fell. Depending on which one did, the fix points in opposite directions.
References
- [1]Peec AI, analysis of the YouTube snippet ChatGPT retrieves, based on the reported ChatGPT retrieval leak, 2026
- [2]Abondance, "YouTube, Facebook, Wikipedia, Forbes : ChatGPT cite beaucoup moins ces domaines", 2026-09-28
- [3]Tomek Rudzki (Peec AI), "ChatGPT built its own search index", Peec AI Blog, 2026
- [4]YouTube Help, "Video chapters"
Summary
- In the published example, the YouTube result ChatGPT received was a text snippet with the title, channel, publish date, views, likes and description, and no transcript.
- The description was longer than all other fields combined, so most of the text the model had about the video was the description.
- In the same example, the YouTube page showed 64,410 views while the snippet showed 56,729, an 11.9% gap. That signals the snippet is a cached document.
- The probability of a citation is P(cite | retrieved) times P(retrieved), so if you don't know which term drove YouTube's 91% drop in citation share, the fixes diverge.
- Timecodes and lists survive into the snippet intact and links become reference tokens, so structuring the description like a landing page pays off either way.
Check this topic against your own brand
TRAIL Search measures how ChatGPT and Perplexity answer your customers' questions and finds and fixes where the brand is missing, turning this post into a diagnosis. Start with 10 questions, no card required.
More posts

How is AI answer visibility defined and measured?
How is AI answer visibility defined and measured? Three metrics with formulas and paper figures: visibility (GEO), contribution (Shapley), and stability.

The prompt you track is not the query AI actually searched
AI visibility reports score each tracked prompt, but retrieval and citation happen on fan-out sub-queries. What year and English injection mean for measurement.

Which content types AI cites, by search intent
The content type AI search cites flips with intent: articles lead informational prompts (45.5%), listicles lead commercial ones (40.9%).
Keep reading on this topic
- How are retrieval and citation different in AI search?Retrieval and citation are different stages in AI search. Does ChatGPT cite YouTube? Why our content is missing from AI answers, and what zero citations mean.
- The LLMO trap: training data vs. AI answer citationsSeeding your brand in LLM training data and getting cited in AI answers are different mechanisms. How they differ, what you can measure, and why advice splits.
- What marketers worry about most in AI search: reporting dataMarketers worry less about vanishing from AI search than about lacking reliable reporting. We checked whether that differs by company size.
This post is part of the Analysis category, which collects all 34 posts on the topic. See all posts in Analysis