Adding a diversity reward to query fan-out: Google's R4T

Adding a diversity reward to query fan-out: Google's R4T

The layer that expands queries is being trained to cut redundancy. Whether that diversity is right for search is a separate question

R4T trains query fan-out on groundedness, diversity, and alignment rewards, then distills it into a small diffusion model. Is that diversity right for search? R4T, from researchers at Google Research and UIUC, trains a fan-out language model with reinforcement learning, uses its outputs to synthesize training data, and then moves that behavior into a small diffusion model.

By · TRAIL Labs Research
query fan-outGEOAEOinformation retrievaldiversity rewarddiffusion models

Google Research's R4T paper trains query fan-out on set-level rewards for groundedness, diversity, and alignment with the original query, then transfers that behavior into a 53.9M-parameter diffusion model that produces a whole set of queries at once. It is a signal that the layer that expands queries is being explicitly trained to reduce redundancy, and the paper does not show whether that diversity is also right for factual search. The paper is "Efficient, Property-Aligned Fan-Out Retrieval via RL-Compiled Diffusion" by Pengcheng Jiang, Judith Yue Li, and nine other researchers from Google Research and the University of Illinois Urbana-Champaign (UIUC) [1]. It was presented as an ICML 2026 poster [3], and Google Research published an explainer on its blog in September 2026 [2].

First, the scope of evidence and how we evaluated it. As of October 2026 we compared two public sources, the arXiv 2603.06397 v1 paper (submitted 2026-03-06) and the Google Research blog explainer, and every number here comes from the paper's text and tables [1][2]. We did not reproduce the experiments ourselves, and the reported speedups are latency results on set-valued recommendation tasks. The paper reports no results on web search rankings or AI answer citations, so please read the practical interpretation in this post with that gap in mind. We have no relationship with the authors, and we note the bias of reading this as a team that runs an AI search visibility diagnosis product.

Why query fan-out is a set problem

Whether fan-out retrieval is good or bad is decided by the whole set, not by any single sub-query. The paper calls this a "set-valued and non-decomposable" retrieval problem [1]. Given a broad intent, the system has to return a collection instead of one result, and that collection has to satisfy higher-order properties such as diversity, intent coverage, and complementarity. At the same time, it has to stay grounded in items that actually exist in the database.

The Google Research blog explains the problem with a camping gear example. Someone who searches for "camping gear" does not want ten slight variations of four-person tents [2]. A good answer is a complementary set: a tent, a sleeping bag, a stove, a headlamp.

Comparison of a failure mode that repeats near-identical queries with a complementary query set, plus default weights for the three reward terms and latency by batch size

Figure 1. On the left, paraphrastic collapse, which produces sub-queries with nearly the same meaning. On the right, a query set spread out so the parts don't overlap. The bottom strip summarizes the default weights of the three reward terms and the latency at batch sizes 8 and 1024. The example queries are illustrative.

This exposes the limit of existing training approaches. Supervised data usually comes as (query, single correct answer) pairs, so it only teaches top-1 retrieval [1]. A condition like "these queries should complement each other" cannot be written as a label on an individual query. There is not a single correct set either. Several different sets can fit the same broad intent, which makes collecting ground-truth labels expensive and subjective.

Reinforcement learning fits this problem because it can optimize a set-level reward directly. The catch is that running an RL-tuned language model at inference time on every request is too slow. It generates sub-queries one after another and calls retrieval for each one. The paper's starting point is to pull those two requirements apart.

R4T's three stages: RL once, inference with a diffusion model

R4T (Retrieve-for-Train) uses reinforcement learning only once, as an "objective transducer," instead of as the deployed inference engine [1]. It does the expensive exploration offline, one time, and distills the result into a small model. That is what "RL-Compiled" in the title means.

Diagram of R4T: stage 1 trains a fan-out language model with RL, stage 2 synthesizes training data, stage 3 trains a diffusion retriever

Figure 2. The three stages of R4T. In stage 1, the fan-out language model (FOLM) generates sub-queries and receives an RL reward from a property check on the retrieved results. In stage 2, the trained FOLM synthesizes query and target pairs, and in stage 3, a diffusion retriever learns to take a query embedding and output a set of target embeddings [1].

StageWhat it doesSetting in the paper
1. Train the fan-out language modelGenerate k sub-queries from a broad query, retrieve with a frozen retriever, apply RL with a set-level rewardGemma3-4B and Qwen3-4B, GRPO with Soft-PPO regularization, group size 8
2. Synthesize training dataRun the trained FOLM, collect high-reward trajectories, convert them into (query embedding, set of target embeddings) pairs128 samples per query, temperature 0.9
3. Train the diffusion retrieverGenerate the whole set of target embeddings at once, conditioned on the query embedding6-layer, 16-head transformer, 128-dim embeddings, set length 12, 256 sampling steps

Table 1. R4T's three stages and the main settings reported in the paper [1].

There are two deployment variants. R4T-FOLM uses the RL-trained language model directly at inference time, and R4T-Diffusion uses the diffusion model that distills its behavior [1]. Comparing the two separates "the quality of reward-optimized fan-out" from "the cost of deploying it." When building the diffusion model's training targets, the rows inside each set are randomly shuffled. A set has no order, so this keeps the model from depending on one.

The three reward terms: groundedness, diversity, alignment

R4T's reward is a weighted sum of three terms, groundedness, diversity, and alignment with the original query, with default weights of 0.6, 0.2, and 0.2 [1]. For the open-ended abstract retrieval (OAR) task, which has no ground-truth set, the reward is as follows.

is the original broad query and is the generated set of sub-queries. Groundedness measures how close each sub-query embedding is to its nearest item in the database. Queries that point at things that actually exist score higher. Alignment is the average cosine similarity between each sub-query and the original query, so it penalizes queries that drift off the question.

Diversity takes one representative retrieved result per sub-query (, for example the top-1 item) and computes the Vendi Score over that set of embeddings [1][4]. The Vendi Score is the exponential of the Shannon entropy of the eigenvalues of a similarity matrix, so it reads as a count of how many effectively distinct things are in the set [4]. What matters is that it measures the diversity of the retrieved results rather than of the query strings. If reworded queries bring back the same items, the diversity score does not go up.

Concept diagram of the three rewards: groundedness pulls sub-queries toward real items, diversity pushes representative results apart, alignment pulls toward the original query embedding

Figure 3. The three reward terms for the OAR task. Groundedness pulls sub-query embeddings toward real database items, diversity (Vendi) pushes representative results apart, and alignment pulls each sub-query toward the original query embedding [1].

In the weakly supervised compositional retrieval (WSCR) task, where part of a reference set is given, the reward is different. It is the share of the reference set that shows up in the fan-out results [1]. The paper treats that reference set as "one plausible realization, not the unique correct answer." So recall on this task is read as a proxy for semantic coverage rather than as accuracy.

Without the diversity term, the policy collapsed

The reason for the diversity term is in the ablation. Trained on the groundedness reward alone, the policy converged to meaningless strings such as "line ending line ending line ending" [1]. It had found strings that happen to sit close, in embedding distance, to a specific database item. With groundedness and alignment together, collapse came even faster. The policy simply repeated paraphrases of the original question and maxed out the alignment score.

The authors call this failure paraphrastic collapse. For the query "Bohemian festival style," the baseline model (Qwen3-4B zero-shot) produced near-identical expressions such as "bohemian festival style" and "bohemian festival fashion," while R4T branched in different directions: "bohemian festival dress," "straw boots festival style," and "lace bohemian festival" [1]. Only by optimizing all three terms together were the shortcuts blocked and training stabilized.

There is also an experiment that changes the weights. When groundedness dominated, diversity rose but alignment fell, and when alignment and diversity were weighted up, exploration shrank and fewer alternative interpretations were covered [1]. The paper concludes that diversity and alignment act as "mutual counter-anchors." In practical terms, diversity in this system is tuned to spread out only within the bounds of the original question, not without limit.

Results: higher quality scores, 12–20x lower latency

R4T scored higher than zero-shot fan-out and the Best-of-N baseline on two datasets and two base models [1]. The datasets are the Polyvore fashion benchmark, built from user-curated outfits, and a proprietary industrial dataset of expert-made music playlists [1][5]. The candidate pools held 21,888 items for Polyvore task 1, 142,472 for task 2, and 8,522 for music. Every fan-out method generated 10 sub-queries, and Best-of-N ran zero-shot fan-out 5 times and kept the highest-reward result.

Method (Polyvore, Gemma3-4B)GroundednessDiversityAlignmentAverage
No fan-out22.434.421.426.1
Zero-shot fan-out28.456.031.238.5
Best-of-N28.961.032.740.9
R4T-FOLM30.876.839.849.1
R4T-DiffusionN/A74.337.6N/A

Table 2. Part of the Polyvore results for the open-ended abstract retrieval (OAR) task. Scores come from an LLM judge, and groundedness was not measured for the diffusion variant because it has no intermediate sub-queries [1].

On the same task, the Gemma3-4B average on the music dataset was 48.1 for zero-shot, 49.2 for Best-of-N, and 58.1 for R4T-FOLM [1]. With Qwen3-4B, the Polyvore averages were 28.1, 30.4, and 42.6 in the same order. The no-fan-out baseline was the lowest in every condition. The paper sums this up as "fan-out is necessary but not sufficient." The long-standing hypothesis that query expansion covers more semantic facets held up, and zero-shot fan-out showed its weaknesses: drifting off the database or collapsing into near-identical queries.

The weakly supervised compositional retrieval (WSCR) task revealed a trade-off between coverage and diversity. The Qwen-based R4T-FOLM had the highest coverage, with Recall@5K of 20.9 and Hit@5K of 64.6, but its Vendi Score of 27.5 was lower than zero-shot Qwen's 46.4 [1]. The diffusion variant sat in between, at Recall@5K 16.5, Hit@5K 57.5, and Vendi 34.7. The more the reward favored overlap with the reference set, the more the results concentrated on a few dominant meanings. The paper adds that because the reference set is not the unique correct answer, lower recall does not necessarily mean lower quality.

Batch sizeAutoregressive LM fan-outDiffusion model fan-out
8About 1.46 s0.07 s
1024About 50 s4.21 s

Table 3. Latency to generate fan-out for 10 sub-queries. The diffusion model has 53.9M parameters, and the paper summarizes the result as a 12–20x speedup across batch sizes [1].

The speed numbers need a careful read. The diffusion model stayed under one second at small batch sizes, but at batch size 1024 it also took 4.21 seconds [1]. "From 50 seconds to under one second" overstates it. What the paper actually shows is "single-digit seconds at the same batch size, 12–20x faster."

Why diffusion: a set has no order

The reason for choosing a diffusion model lies in the shape of the output. Autoregressive generation emits sub-queries one at a time, in sequence. Nothing says whether the tent or the sleeping bag comes first, yet sequential generation invents an order and outputs in that order. A diffusion model starts from noise, refines the whole set at once, and outputs it in one pass, which matches an unordered set as the output structure [1].

The paper describes this as a shift from heavy "System 2" autoregressive generation to lightweight "System 1" parallel sampling [1]. The diffusion model receives the query embedding through cross-attention, solves a probability flow differential equation to produce a set of embeddings, and then maps each embedding to a real item with nearest-neighbor retrieval. The 12–20x speedup reads as a consequence of this design. The goal was "keeping set-level properties within production latency."

Limitations: it does not show that diversity is right for search

The paper names four limitations itself [1]. First, the RL stage needs repeated interaction with a frozen retriever, so upfront training costs grow for very large or frequently changing databases. Second, it assumes the desired properties can be written as a scalar reward. Subjective preferences such as creativity or cultural sensitivity are hard to turn into rewards. Third, quality on the open-ended task was judged by an LLM, so the judge model's biases can creep in. Fourth, the experiments used one fan-out architecture and one diffusion retriever, so results may differ with other base models or embedding spaces.

To us, the bigger limitation is the evaluation domain. In fashion outfits and music playlists, complementarity goes beyond taste to become a condition of a correct answer. An outfit with four tops is not a boring outfit; it is a wrong one. So putting diversity into the reward is a natural design in those two domains. Whether web search is that kind of domain has to be shown separately, and this paper does not make that claim.

Factual questions make the difference clear. Three trustworthy documents that say the same thing are better grounds for an answer than ten documents that each say something different. There, redundancy is corroboration, and a diversity reward works to erode that corroboration. From these results alone, we cannot tell whether the premise that diversity is good rode in on the evaluation domain or holds for search in general.

There are also details in the write-up worth checking. The main text names the LLM judge as Gemini-2.5-Pro, while the appendix says Gemini-2.5-Flash [1]. The music dataset is proprietary and cannot be reproduced externally, and the conclusion paragraph mentions only the fashion benchmark. The diversity and groundedness scores were also assigned by an LLM judge from the same family.

What it means for AEO/GEO: is a good answer in your category complementary or corroborating?

The conclusion to take from this paper is narrower than the headline. We cannot say "engines reward diversity." The well-supported conclusion stops here: in domains where a correct answer must be a set of complementary parts, the retrieval layer is being built to enforce that explicitly, and it has become light enough to fit within production latency.

Still, one direction is clear. When fan-out is optimized at the set level, a single page is not matched against a single question. It is matched against a set of queries deliberately spread out so they don't overlap, and it either wins one slot in that set or it doesn't. A page that answers only the most common phrasing, however well written, gets one slot at most. For how fan-out shows up in real AI search, see What is query fan-out? Watch AI split one question, and for how to reflect sub-queries in titles and H2s, see Query fan-out GEO: 5 steps to put subqueries in titles.

This turns a familiar content question into one you can test.

  1. First decide what a good answer looks like in your category. In some categories, such as camping gear recommendations, a good answer is a set of complementary sources. In others, such as fact checks, a good answer is a set of sources that agree.
  2. In a complementary category, keeping many similar pages becomes a liability. Adding more pages on the same facet still competes for only one slot in a set spread out for diversity. One page per facet fits the set better.
  3. In a corroborating category, redundant coverage becomes a hedge. When several sources confirm the same fact, consistent information in several places acts as a corroboration signal.
  4. Measure instead of assuming. You can observe whether the sources an AI answer cites for the same question cover different facets or repeat the same content.

For the landscape of tools and measurement approaches around fan-out, continue with What is AI search query fan-out? Find hub queries, adapt. What this paper shows is the design direction of the retrieval layer. Which way that design works in your category remains a question each team has to measure.

Frequently asked questions

What does R4T do?

It is a three-stage method. It first trains query fan-out, which expands one broad question into several sub-queries, with reinforcement learning, then transfers that behavior into a 53.9M-parameter diffusion model. At inference time the diffusion model, not a language model, produces several retrieval directions at once from a single query embedding.

Why was a diversity reward needed?

In the paper's experiments, a policy trained on the groundedness reward alone collapsed into meaningless strings, and one trained on groundedness plus alignment just repeated paraphrases of the original query. Only with the diversity term added were these shortcuts blocked and training stable.

Can these results be applied directly to web search or AI answer citations?

Not yet. The evaluation tasks were fashion outfits and music playlists, domains where complementarity is part of what makes an answer correct. For factual questions, several trustworthy documents that say the same thing act as corroboration, and the paper reports no results on web search or citations.

How much faster did it actually get?

Per Figure 5 of the paper, at batch size 8 the autoregressive approach took about 1.46 seconds versus 0.07 seconds for the diffusion model, and at batch size 1024 about 50 seconds versus 4.21 seconds. The paper summarizes this as a 12–20x speedup. It did not drop below one second at large batch sizes.

References

  1. [1]Jiang, Li, Ryu et al., "Efficient, Property-Aligned Fan-Out Retrieval via RL-Compiled Diffusion", arXiv 2603.06397 (ICML 2026)
  2. [2]Google Research Blog, "Bypassing inference bottlenecks: Accelerating complex AI search with Retrieve-for-Train" (2026-09-15)
  3. [3]ICML 2026 poster session page, "Efficient, Property-Aligned Fan-Out Retrieval via RL-Compiled Diffusion"
  4. [4]Friedman & Dieng, "The Vendi Score: A Diversity Evaluation Metric for Machine Learning", arXiv 2022
  5. [5]Han, Wu, Jiang & Davis, "Learning Fashion Compatibility with Bidirectional LSTMs", ACM MM 2017 (Polyvore dataset)

Summary

  • R4T, from researchers at Google Research and UIUC, trains a fan-out language model with reinforcement learning, uses its outputs to synthesize training data, and then moves that behavior into a small diffusion model.
  • The reward has three terms, groundedness, diversity, and alignment with the original query, with default weights of 0.6, 0.2, and 0.2. Diversity is measured with the Vendi Score.
  • The design rests on an ablation: without the diversity term, the policy collapsed into meaningless strings or paraphrases of the original query.
  • The diffusion model produces retrieval directions for 10 sub-queries in one pass, ran 12–20x faster depending on batch size, and scored higher than the baselines on quality.
  • Because the evaluation domains are outfits and playlists, this paper cannot tell us whether a diversity reward is right for factual web search.

Check this topic against your own brand

TRAIL Search measures how ChatGPT and Perplexity answer your customers' questions and finds and fixes where the brand is missing, turning this post into a diagnosis. Start with 10 questions, no card required.

The feature closest to this post is Keyword and question discovery.

More posts

Keep reading on this topic

This post is part of the Papers category, which collects all 15 posts on the topic. See all posts in Papers