Does putting the year in the title increase AI citations?

Does putting the year in the title increase AI citations?

A year-in-title effect and a recency effect cannot be separated by observation alone

The 170% claim mixes a string effect with a recency effect. How much recency matters for AI search citations, and how to test the effect with a 2×2 experiment.

By · TRAIL Labs Research
GEOAEOAI citationrecencyexperiment designcausal inference

The conference claim that adding this year to a title lifts AI citations sharply merges two things: the effect of the year string itself and the effect of the page actually being recent. Following it blindly, you cannot know which one you moved. The "2026" in a title is also a signal that the page is from this year, so in observational data the year label and recency move almost perfectly together. Answering the question takes an experiment, not an observation, and the design is simpler than it sounds.

First, the scope of our evidence. The 170% figure is circulating in conference notes, and we could not check its underlying data or methodology, nor did we reproduce it. So we are not saying it is wrong. The weight of recency comes straight from the tables of the Sprinklr team's 252,000-run controlled experiment (What Gets Cited, SIGIR 2026) [1], and we cross-checked the need for repeated measurement against the Don't Measure Once paper [2]. The fan-out observation is primary data from a vendor researcher's LinkedIn posts with no peer review, so by our standard its evidence grade is medium or below and we do not use it to justify any weighting [6]. This reflects what we knew as of October 2026, and since one passage mentions how our own diagnosis product scores pages, we disclose that interest too.

The 170% year-in-title figure is the sum of two effects

The observed difference splits into two terms. If is the citation rate difference actually seen between pages with the year in the title and pages without it, you can write:

is the effect of the string in the title itself, the share that changes when you add only the year to the same page. is the share created by the recency signal the year carries, and by the fact that pages with a year in the title were actually written or updated recently. Statistics calls this a decomposition into direct and indirect effects, and in linear models the total effect equals their sum [5]. The full definitions are in Wikipedia's entry on mediation).

The problem is that observational data shows you only the sum. Pages with this year in the title were usually written this year or recently revised. Because the label and actual recency move together, whatever the gap between the two groups turns out to be, there is no way to compute how much of it belongs to the string. is not identified.

Confound diagram showing two paths between a year in the title and an AI citation

Figure 1. Actual recency affects both the year label in the title and the AI citation, while the direct path from the year label to the citation cannot be identified from observation alone. The box at the bottom shows that the observed difference is the sum of the two terms.

How much does recency matter for AI search citations?

Recency is a gatekeeper-level factor in AI citation, so the weight of the recency path is not small. The What Gets Cited paper attached 18 content factors one at a time to six commercial LLMs, asked 252,000 times, and classified four of them as gatekeepers [1]: topic match, a stated price, a recent timestamp, and a front slot in the context. The paper concludes that if any one of the four fails, citation odds vanish regardless of other strengths.

Here are the per-model odds ratios for the three recency factors as reported. An odds ratio tells you how many times higher the odds of being cited are for the favored side (listed first) than for the weaker side.

Factor (favored vs. weaker)Gemini-2.5-FlashClaude-3.5-SonnetKimi-K2GPT-5-NanoGPT-5-MiniGPT-5.2
Recent (2026) vs. old (2019) dateover 10kover 10k68.714.41,494over 10k
No date vs. old date2.322.281.331.311.551.48
Recent date vs. no date1.9913.03.671.153.531.99

Table 1. Odds ratios for the recency factors in What Gets Cited, under the paper's experimental conditions [1]. The first row was significant in all six models; the two rows below were significant in only some of them.

Two things stand out. First, pitting a recent date against an old one produces a very large effect. Second, the ranking settles as recent > no date > old. Having no date at all generally did better than an old date. In this experiment the recency treatment kept facts, price, specs, and length fixed and changed only the date in the body to 2026 or 2019 [1]. Between two candidates already in the context, the date label alone moves the choice a lot, which is exactly why you cannot ignore the size of and read the conference lift as a string effect. We covered the full design and the other gatekeepers in Which of two pages will AI cite?.

How the controlled experiment isolated one factor

What to learn from What Gets Cited is the design more than the results. The researchers picked 100 real review blogs from 50 B2C categories and replaced brand and publisher names with fictional aliases [1]. The point was to remove the bias toward names a model already knows from pretraining. They then built two variants that differed in exactly one factor, presented the two candidates in both orders, and repeated with three different query phrasings.

The entire design is spent on isolating one factor. Effect sizes were estimated separately for each factor with a logistic mixed-effects model that includes random intercepts for scenario and order [1], which also corrects for repeats within the same scenario resembling each other.

The lift figure from the conference has none of these controls. It is an observation, not an experiment. That does not make observation bad. The trouble starts when it reads as if it answered a question observation cannot answer: does the year string raise citations?

One more point from the same paper. The "front slot" gatekeeper is a different construct from putting key content at the top of your page. It is the slot order after a page has been retrieved into the answer context. And only after the four gatekeepers passed did seven secondary factors decide the winner, with odds ratios mostly between 2.1 and 243 [1]. Dense paragraphs versus organized sections, on the other hand, showed odds ratios of 0.78–1.68 with no consistent direction across the six models. There is an order of operations. Tuning secondary factors before clearing the four gatekeepers is polishing a stage you have already been cut from.

There is also a path where the year string acts directly

There is no basis for assuming is zero either, because a mechanism through which the string could act directly has been observed. A vendor researcher turned Google News trending stories into questions and automatically asked ChatGPT more than 8,000 of them over a month. Only 1% of the original prompts contained a year, yet "2026" appeared in 94% of the follow-up searches the engine generated (fan-out queries) [6]. The original was shared in Malte Landwehr's LinkedIn posts.

If this observation holds, then the moment the engine appends a year to its own searches, a document with the same year in its title matches better on vocabulary at the retrieval stage. That is string matching unrelated to recency, and it belongs squarely in . We covered how fan-out splits a question in our post observing query fan-out.

We want to be clear about the grade of this observation, though. It is a single observation on news topics, in English, centered on ChatGPT, and it was not peer reviewed. So this figure is material for the hypothesis that a string path exists, not evidence for its size. The fact that both paths are plausible is exactly why an experiment is needed.

Engines differ in sensitivity, so a single average is the wrong number

A single rate pooled across engines overstates the effect for some engines and understates it for others. The first row of Table 1 alone shows odds ratios for the same recency treatment ranging from 14.4 for GPT-5-Nano to over 10,000 for Gemini and Claude [1]. The paper reports that all six models agreed on the four gatekeepers but differed in sensitivity below them. Kimi-K2 responded significantly to 83% of the 18 factors and Gemini-2.5 to only 33%, but when Gemini did respond it was extreme, like a binary switch.

So a single lift figure like the conference number that does not separate engines erases that variance. It becomes an actionable number only when it comes with which engine it was measured on and how it breaks down by engine. We covered the same problem from the reporting side in Why scoring AI visibility with one number is risky.

How to test AI citation effects with an experiment: a 2×2 design

The way to separate the two terms is to move one factor at a time. Run pairs that keep the content and change only the title year alongside pairs that keep the title and genuinely refresh the content, and you fill four cells that separate the two terms.

CellTitle yearContentWhat this cell shows
ANoneUnchangedBaseline
BAdd this yearUnchangedEstimate of the string effect (B − A)
CNoneGenuinely refreshedEstimate of the recency effect (C − A)
DAdd this yearGenuinely refreshedWhether the two effects add up or interact (D − B − C + A)

Table 2. A 2×2 design that splits the year-in-title effect into a string path and a recency path. This is the design TRAIL Labs proposed in our LinkedIn post, laid out as a table.

One condition must always come with it. Ask the same query again on the same day and the cited sources change quite a bit. The Don't Measure Once study reported that even runs repeated immediately on the same day overlapped only 0.32–0.43 in cited sources (Jaccard), essentially the same range as day-to-day comparisons (0.34–0.42) [2]. Much of the variance comes from the model's probabilistic generation rather than the passage of time. In the same study, bringing the standard error of brand detection below 0.10 took 7 repeated runs [2]. All of these are under the paper's experimental conditions.

So with a single read per cell, you cannot tell whether the differences among the four numbers are treatment effects or that day's variance. Measure each cell several times and report how many runs produced each value alongside the results. For setting a baseline and choosing a measurement cadence, you can use the procedure in How to design AI search visibility experiments as is.

Does changing only the content date help in AI search?

Changing only the content date does not help, and becomes risky, the more the recency path dominates. Depending on which term is larger, the recommendation flips. If the direct effect dominates, you only need to change the title string and the cost is close to zero. If the recency path dominates, changing only the title does nothing, and making stale content look new by updating the year is actually risky.

The reason we call it risky lies in search guidelines. Google lists "Are you changing the date of pages to make them seem fresh when the content has not substantially changed?" as a warning sign of search engine-first content [3]. The original is in Google Search Central's guide to helpful content. Its guidance on dates also says a date must describe the publication or update date of the page [4], as stated in the byline date guide.

In our diagnosis scoring, freshness accounts for 11 of the 32 authority points. That is a meaningful share, and it is why we designed the score to look at whether content was actually updated instead of at the date string. If a score can rise just because a label changed, it is better that it does not rise.

Limitations

Here is what this post did not see. We could not check the underlying data or measurement method behind the conference lift figure, and we have not run that figure or the 2×2 test ourselves. Table 2 is a proposed design, not a result.

The What Gets Cited results measure the relative preference between two candidates already in the context [1]. The paper does not cover which documents get into the candidate set at the retrieval stage, and it did not test putting a year in the title as a separate treatment. The paper's text summarizes the gatekeepers as very strong factors with odds ratios above 100, but as Table 1 shows, some models' values are smaller. Its setting, 100 B2C review blogs with anonymized brands, also does not transfer directly to other industries or to Korean-language queries.

The 94% fan-out observation comes from a vendor researcher's social media posts with no peer review, and it is a single observation on news topics, in English, centered on ChatGPT [6]. Part of the sample and method was disclosed, but we did not reproduce it, and since the value can shift when model versions change, it should not be read as a constant. The variance figures in Don't Measure Once come from four Swiss German-language campaigns across four engines [2].

Next steps

If you are told that "putting the year in the title lifts citations," first ask whether there is a way to check if that comes from the string or from recency. If there is none, editing titles in bulk based on that figure is a decision made without knowing which term it touches. The practical order is to refresh the pages worth refreshing first, label update dates honestly, and test the title year separately, as in cell B of the 2×2.

Frequently asked questions

Does adding this year to a title really increase AI citations?

Even if you observe a lift, observational data cannot tell you whether it came from the year string or from the fact that pages with a year in the title are actually recent. You need to measure pairs where only the title year changes and pairs where only the content is refreshed.

How much does recency matter for AI citations?

In a 252,000-run controlled experiment that pitted two candidates against each other, a recent date versus an old date was significant in all six models, and the paper classified it as one of four gatekeepers alongside topic match, a stated price, and a front slot. Odds ratios ranged from 14.4 to over 10,000 by model. All figures are under the paper's experimental conditions.

Can I keep the content as is and just update the date?

We do not recommend it. Google lists changing a page's date to make it seem fresh when the content has not substantially changed as a sign of search engine-first content. If the effect runs through recency, changing only the label does not touch the cause and only costs you trust.

How many times should each cell of the test be measured?

Research shows that even when you ask the same question again on the same day, only about 34–42% of cited sources overlap. With one read per cell you cannot tell a treatment effect from that day's variance, so measure each cell several times and report how many runs produced each number.

References

  1. [1]Vishwakarma, Kumar & Jamidar, "What Gets Cited: Competitive GEO in AI Answer Engines", SIGIR 2026
  2. [2]Schulte, Bleeker & Kaufmann, "Don't Measure Once: Measuring Visibility in AI Search (GEO)", arXiv 2026
  3. [3]Google Search Central, "Creating Helpful, Reliable, People-First Content"
  4. [4]Google Search Central, "Add a Byline Date to Google Search Results"
  5. [5]Wikipedia, "Mediation (statistics)": definitions of direct, indirect, and total effects
  6. [6]Malte Landwehr (Peec AI), LinkedIn posts: automated observation of 8,000+ news prompts (2026, not peer reviewed)

Summary

  • The claim that a year in the title lifts citations 170% is the sum of an effect from the string itself and an effect from the page actually being recent. Observation shows only the sum, and neither term is identified.
  • Recency is not a minor factor. In a 252,000-run controlled experiment, a recent date was one of four gatekeepers, with odds ratios ranging from 14.4 to over 10,000 by model (under the paper's experimental conditions).
  • There is also a path where the string acts directly. In one observation of news queries, only 1% of original prompts contained a year, but 94% of the engine's follow-up searches did. It is a social media observation, so its evidence grade is medium or below.
  • A 2×2 test separates the two paths. Run pairs that change only the title year alongside pairs that refresh only the content, and measure each cell several times.
  • If the direct effect is large, editing titles is cheap and correct. If the recency path dominates, putting a new year on stale content brings risk with no effect.

Check this topic against your own brand

TRAIL Search measures how ChatGPT and Perplexity answer your customers' questions and finds and fixes where the brand is missing, turning this post into a diagnosis. Start with 10 questions, no card required.

More posts

Keep reading on this topic

This post is part of the Analysis category, which collects all 34 posts on the topic. See all posts in Analysis