
The prompt you track is not the query AI actually searched
A prompt tracking report's score attaches not to one question but to the set of queries the engine expands it into
AI visibility reports score each tracked prompt, but retrieval and citation happen on fan-out sub-queries. What year and English injection mean for measurement.
The per-question score in an AI search visibility report is not the score of that one question string; it is the result of the set of sub-queries the engine expanded the question into through query fan-out. The engine searches with sub-queries that add a year or switch to English, and how much each branch contributed to the answer is not visible from outside. So if you read the scores you get for 30 tracked questions as scores for those 30 strings alone, you can attribute the results to the wrong cause.
First, the scope and grade of our evidence. As of October 2026 we compared five public sources: one social post, three vendor analyses, and one industry article covering one of those analyses [1][2][3][4][5]. All of them are observations that were not peer reviewed, so on our evidence scale they rank medium or below, and we did not reproduce any of the numbers. The vendor analyses come from a company that sells AI visibility tracking, and we also run an AI search visibility diagnosis product, so there are interests and bias from the same market on both sides. This post uses these numbers only to set the direction of measurement design, not as grounds for scoring weights.
The prompt you track is not the query the engine searched
The tracking unit is the question, but the retrieval unit is the sub-query. A visibility report usually works by entering questions and getting scores per engine, but the engine does not search the question as-is. It expands it into several branches. That process is query fan-out, and we covered how to observe it directly in What is query fan-out? Watch AI split one question.
How far and in what shape the question expands is not fixed. In the Peec AI blog's analysis of 5 million fan-outs, ChatGPT generated 2.1 sub-queries per prompt on average, Perplexity 1.4, and Grok 6.8 [4]. In a separate analysis from the same company, the average length of ChatGPT sub-queries doubled in four months, from about 6 words in October 2025 to about 12 words in January 2026. The number of sub-queries per prompt stayed nearly the same over that period, at 2.3–2.8 [5]. Even for the same question, the strings the engine actually searches change over time.

Figure 1. The tracked question q passes through a fan-out F that cannot be seen from outside and expands into several sub-queries, and retrieval and citation happen on those branches. The figures on the right come from separate public observations, and the example queries are illustrative.
Two public observations: year injection and English sub-searches
There are two lines of public observation showing that fan-out adds strings the original question did not contain. One is the year, the other is language.
| Observation | Sample and method | Figure | Evidence grade |
|---|---|---|---|
| Year injection on news prompts | Google News trending stories turned into questions and re-asked every 6 hours, 8,000+ in one month, ChatGPT | 94% of follow-up queries contained "2026," vs. 1% of original prompts | Social post, medium or below |
| Year injection on general prompts | 5 million fan-outs, April 1–21, 2026 | ChatGPT added the current year in 5.44% of prompts | Vendor blog, medium or below |
| English sub-searches on non-English prompts | February 2026, 10M+ prompts and 20M fan-outs, only cases where IP location matched the question language | 78% of non-English sessions had at least one English sub-search; 43% of all sub-queries were in English | Vendor research, medium or below |
Table 1. Public observations showing that fan-out adds strings the original question did not contain [1][2][4].
The year observation comes from a public LinkedIn post by Malte Landwehr [1]. It covers more than 8,000 automated conversations built from trending stories on the Google News home and topic pages, re-asked every 6 hours until the 24-hour news cycle ended. For news prompts, ChatGPT ran a web search 100% of the time, and 94% of the follow-up queries contained "2026." Only 1% of the original prompts included a year.
This number has to be read by prompt type. In an analysis of 5 million general customer prompts, ChatGPT added the current year 5.44% of the time [4]. The picture is that for questions where timing is part of the answer, like news, year injection is close to the default, and for other questions it is rare.
The English observation is an analysis Peec AI Research published in February 2026 [2]. 78% of non-English sessions included at least one English sub-search, and by language the share ranged from a low of 66% for Spanish to a high of 94% for Turkish. The Search Engine Journal article covering this result noted that prompt composition and representativeness were not disclosed, and that the data came from the vendor's measurement platform rather than from consumer sessions [3]. Neither source gives a figure for Korean.
Beyond years and languages, the engine inserts words
The strings added to sub-queries are not limited to years and languages. In the analysis of 5 million fan-outs, the words added most often were "best," "top," "comparison," "reviews," "tools," and "software" [4]. For advice-seeking questions, "best" appeared in 24.3% of sub-queries. Even if a user only asks for "AI visibility measurement tools," the string the engine actually searches may look closer to "best AI visibility tools comparison."
Sub-queries that target specific sources were observed too. In the same analysis, Grok used Reddit in its searches in 10.5% of all conversations, and 9 out of 10 of those were searches that named Reddit explicitly with the site: operator [4]. In cases like this, the candidate document pool is assembled within a source range the engine chose, not from the string the user typed.
| What gets inserted | Observed example | Effect on measurement |
|---|---|---|
| Timing | "2026" in 94% of follow-up queries for news questions | Whether a page's year matches the engine's string gets mixed into the score |
| Language | English sub-searches in 78% of non-English sessions | The English document pool's share gets mixed into Korean scores |
| Modifiers | "best" in 24.3% of sub-queries for advice questions | Builds pools that favor comparison and recommendation pages |
| Source targeting | Reddit searches in 10.5% of Grok conversations, mostly via site: | Creates slots where you only compete inside one platform |
Table 2. Types of strings inserted into sub-queries and their effect on measurement. The figures come from separate observations [1][2][4].
None of these four appear in the question string in your report. Even with the same tracked prompts, when the strings the engine inserts change, the competing document pool changes, and the score moves with it.
A report score is V(F(q)), not V(q)
The value printed as the score for question q is really a weighted sum of sub-query scores. If we call the tracked question and write that the engine applies an invisible transformation to produce a set of queries, it looks like this.
is how well our page is retrieved and cited for one sub-query, and is that branch's contribution to the final answer. Retrieval and citation happen on , not on . The problem is that neither nor can be observed from outside. What we measure is wearing the name .
Two consequences follow. One is that attribution can go wrong, and the other is that the language boundary is thinner than it looks. The next two sections take them in turn.
The attribution of the year effect changes
The year injection observation adds a candidate mechanism to the claim that "putting this year in your title raises citations." We covered how that number cannot separate a freshness signal from the effect of the title string itself in Does putting the year in the title increase AI citations?. The 94% observation adds a third explanation to that gap [1]. If users rarely type a year but the engine adds one to its sub-queries, then a year in your title is matching a string the engine inserted rather than one the user typed.
This is not proof. It is a different observation, on news topics, and correlational. What it does give us is testable predictions.
- The year effect should be larger for question types that get more year injection. It should be large for news and latest-comparison questions and small for definition questions.
- It should bend once when the year turns over. When the token the engine injects changes from "2026" to "2027," the effect for pages with "2026" in the title should drop at that point.
- If it doesn't bend, the string-matching explanation weakens. In that case the freshness-signal explanation remains.
In practice, whether to put a year in a title is a call to make by question type. If you stamp a year on a post you won't update, you also take on the risk of mismatching the string the engine injects once the year turns over.
English branches get mixed into scores tracked in Korean
A score tracked with Korean questions is likely not the result of Korean searches alone. When one sub-query goes out in English, part of the candidate document pool is assembled from English documents. A brand that targets only the domestic market is competing for only some of the slots to begin with.
Here is the counterargument too. One sub-query being in English does not mean the final answer is filled with English sources. The figure that 43% of sub-queries are in English also means the other 57% stay in the original language [2]. Results vary by query type, by whether the topic is local or general, and by engine, and Korean does not appear as a figure in the observed samples. So the conclusion from this observation does not go as far as "you need to publish English content." It stops at this: an English branch's share may be mixed into your Korean scores, so look at that share separately.
What you can do from outside: measure the shadow, design for sets, record the version
You cannot see fan-out , but you can see the marks leaves on the output. Some surfaces have been reported to expose internal search queries, but that is not a stable observation path. Instead, here are three things you can do now.
| What to do | What to look at | How far you can take it |
|---|---|---|
| Record the language and domain mix of cited sources | Whether Korean questions come back with citations mostly from non-Korean domains | A shadow of the invisible English branch. A proxy, not a measurement of fan-out itself |
| Design prompt sets around the query sets they induce | Whether you collected questions the engine will expand into different sub-queries, rather than just different strings | Tracking scope closer to real coverage |
| Store the engine and model version with every run | Whether score changes line up with model changes | The minimum condition for separating your content's share from the ruler's share |
Table 3. Measurement add-ons you can apply from outside when you cannot see fan-out directly.
First, record the language and domain mix of cited URLs. If a Korean question comes back with citations mostly from non-Korean domains, that is the mark the invisible English branch left on the output. It is a proxy, and it should only be used as one.
Second, change how you build prompt sets. Collecting questions with different strings and collecting questions that induce different sub-query sets are different jobs. The first is easy, and the second is real coverage. "AI search visibility measurement tool" and "AI visibility tracking tool recommendations" differ as strings but are likely to expand into similar sets, so adding both barely increases coverage. We covered the basics of designing monitoring prompts in Prompt tracking: How to monitor AI search citations, and how to reflect sub-queries in your content in Query fan-out GEO: 5 steps to put subqueries in titles.
Third, record on the premise that has versions. The same question expands into different sets over time. In an analysis of more than 20 million fan-outs across five countries, the average sub-query length grew from about 6 words in October 2025 to about 12 words in January 2026, peaking at about 16 words in one week in between [5]. That analysis did not tie the change to a specific model version. Still, when the strings the engine searches change this much while your content stays the same, the candidate pool your page competes in changes too. If the engine version is not stored alongside your measurements, you cannot separate your content's share of a score change from the ruler's share.
What this post doesn't cover
Every number in this post is a public observation, and none were peer reviewed. The 94% year figure comes from a single social post, and since we did not preserve the individual post URL, the citation points to the author's profile [1]. The English sub-search figures come from a vendor platform running customer-defined prompts through browser automation, so there is no basis for treating them as representative of ordinary user sessions [3].
The samples are skewed too. The year observation is limited to news topics, and the language observation includes no figure for Korean. We did not reproduce any of the numbers, and we have not checked whether Korean questions show the same rates. The weighted sum in equation (1) is a model that describes the measurement structure, not a claim that engines actually build answers through this kind of linear combination.
It is safer to read the numbers as a direction than to carry them over as constants. The direction itself is clear. The score you get for a question is less the score of that question than the score of a set that passed through an invisible transformation.
Frequently asked questions
Does AI search run the exact question I typed?
No. AI search engines go through query fan-out, expanding the question they receive into several sub-queries before searching. In public observations, ChatGPT generated a little over 2 sub-queries per prompt on average, and those queries picked up strings the original question did not contain, such as a year or English phrasing.
Does ChatGPT search in English when asked in another language?
According to public vendor observations, often. In a February 2026 analysis, 78% of non-English sessions included at least one English sub-search, ranging by language from 66% for Spanish to 94% for Turkish. No figure for Korean was published.
So are prompt tracking scores meaningless?
They are meaningful. You just need to read them knowing that the score is not for one question string but the sum of results across the set of queries the engine expanded it into. Recording the engine version alongside each run and looking at the language and domain mix of cited sources makes the interpretation less likely to go wrong.
How reliable are these observed numbers?
On our evidence scale, medium or below. They come from social posts and vendors' own analyses, none peer reviewed. The samples also lean toward news topics and English-speaking markets, so it is safer to read the numbers as a direction than to carry them over as constants.
References
- [1]Malte Landwehr, public LinkedIn post: over 8,000 automated ChatGPT conversations in one month on Google News trending topics (fall 2026, collected by TRAIL Labs on 2026-10-04; the individual post URL was not preserved, so this links to the author's profile)
- [2]Peec AI Research (Tomek Rudzki), "ChatGPT searches in English": analysis of 10M+ prompts and 20M fan-outs (February 2026)
- [3]Search Engine Journal (Matt G. Southern), "ChatGPT Search Often Switches To English In Fan-Out Queries: Report" (2026-02-18)
- [4]Peec AI Blog (Tomek Rudzki), "Patterns we see in ChatGPT query fanouts": analysis of 5M fan-outs (2026-10-07)
- [5]Peec AI Blog (Tom Wells), "ChatGPT fan-outs have doubled in length in 4 months": 20M+ fan-outs across 5 countries (2026-02-12)
Summary
- Prompt tracking reports score each question, but the engine expands the question into several sub-queries and retrieves and cites on those.
- One observer reported that 94% of follow-up queries for news prompts contained "2026," while only 1% of the original prompts included a year. In an analysis of 5 million general prompts the figure was 5.44%.
- In a vendor analysis, 78% of non-English sessions included at least one English sub-search. No figure for Korean was published.
- So the V(q) printed in a report is really a sum of sub-query scores weighted by unobserved weights, and it moves when the model changes even if your content does not.
- What you can do from outside is treat the language and domain mix of cited sources as a proxy, design prompt sets around the query sets they induce, and record the engine version with every run.
Check this topic against your own brand
TRAIL Search measures how ChatGPT and Perplexity answer your customers' questions and finds and fixes where the brand is missing, turning this post into a diagnosis. Start with 10 questions, no card required.
The feature closest to this post is Keyword and question discovery.
More posts

How is AI answer visibility defined and measured?
How is AI answer visibility defined and measured? Three metrics with formulas and paper figures: visibility (GEO), contribution (Shapley), and stability.

Which content types AI cites, by search intent
The content type AI search cites flips with intent: articles lead informational prompts (45.5%), listicles lead commercial ones (40.9%).

How are retrieval and citation different in AI search?
Retrieval and citation are different stages in AI search. Does ChatGPT cite YouTube? Why our content is missing from AI answers, and what zero citations mean.
Keep reading on this topic
- The LLMO trap: training data vs. AI answer citationsSeeding your brand in LLM training data and getting cited in AI answers are different mechanisms. How they differ, what you can measure, and why advice splits.
- What marketers worry about most in AI search: reporting dataMarketers worry less about vanishing from AI search than about lacking reliable reporting. We checked whether that differs by company size.
- Should AI visibility be measured through the API or the UI?The same prompts through the ChatGPT API and UI shift brand visibility 41% on average and change cited sources. What to check when choosing a measurement tool.
This post is part of the Analysis category, which collects all 34 posts on the topic. See all posts in Analysis