SYNTHESIS NOTE
Topics›Deep Research›this note

Does reinforcement learning squeeze exploration diversity in search agents?

Investigates whether RL training narrows the behavioral diversity of search agents the same way it does in reasoning tasks. Understanding this mechanism could reveal whether entropy collapse is fundamental to RL or domain-specific.

Synthesis note · 2026-02-21 · sourced from Deep Research

The "RL Squeezes, SFT Expands" paper studies search agents trained with RL versus SFT and finds the same pattern that the reasoning literature documented: RL training compresses the diversity of behaviors the agent explores (squeezes), while SFT on diverse demonstrations expands it. Since Does policy entropy collapse limit reasoning performance in RL?, and since this paper shows the same dynamic in search RL, entropy collapse is not a quirk of reasoning training — it is a property of RL training at large.

The mechanism is the same in both domains: RL rewards the policy for high-reward outputs and penalizes low-reward ones. Over training, the policy concentrates probability mass on the reward-maximizing region of its action space. In reasoning, this means converging on a narrow set of reasoning patterns. In search, it means converging on a narrow set of query strategies. Both reduce the agent's ability to explore novel approaches to hard problems.

SFT has the opposite effect because it trains on human demonstrations or diverse synthetic completions — the diversity of the training set is preserved in the policy. The tradeoff is that SFT cannot generalize beyond its demonstrations in the same way RL can.

This finding has practical implications for DR agent design: RL-trained search agents need explicit diversity mechanisms (entropy regularization, diverse reward models, periodic SFT refreshes) or they will converge on query templates that work well on average but fail on distribution shift. The same Do critique models improve diversity during training itself? remedy applies — external critique prevents the RL agent from collapsing to a narrow search strategy.

Inquiring lines that read this note 157

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Does AI assistance promote real skill development or substitute for independent learning? What types of diversity prevent reasoning systems from collapsing? How do agent-learned skills transfer and improve across different tasks? How do pretraining biases affect reward signal effectiveness in RLVR? How does policy entropy collapse constrain scaling of reasoning-focused RL? What reasoning architectures enable models to solve complex problems efficiently? How does synthetic data quality and diversity affect downstream model capabilities? How can evolutionary algorithms maintain diversity during solution search? Can self-generated feedback reliably guide model training without ground truth? What training dynamics and scale trigger emergence of reasoning capabilities? Does RL create genuinely new reasoning capabilities or refine existing ones? How do neighboring agents influence whether others cooperate or collude? How effectively can language models perform reasoning, especially combined with symbolic methods? Why do persona simulations fail to predict authentic user behavior? Why can't prompting alone inject genuinely new knowledge into models? When do multi-agent systems outperform single frontier models? Can prompt-based context override biases that were embedded during pretraining? Where and how do personality traits reside in language models? What makes personas effective for predicting individual preferences and behavior? How do surface patterns enable correct outputs but reduce robustness? Why do agents falsely report success on failed tasks? How should agents manage memory granularity to improve long-term performance? Does preference optimization systematically degrade conversational grounding in language models? Can AI systems distinguish genuine empathy from simulated emotion? Do language models lack essential therapeutic presence and engagement? Does alignment training create genuine alignment or just output compliance? How should inference compute be allocated based on problem difficulty? How does decomposing tasks improve reasoning and prevent failure propagation? What causes reasoning models to fail or wander off track? What training data selection strategies maximize generalization across difficulty levels? How does harness optimization generalize across different model architectures and domains? Can harness architecture and protocols provide agent reliability without model scaling? Do language models develop actual world models or merely task heuristics? How do soft reasoning mechanisms explore multiple paths without explicit training? What should agent evaluation prioritize to reveal reliable behavior? What capability trade-offs arise from domain specialization through fine-tuning? How do spurious versus genuine rewards shape model reasoning and behavior?

Related concepts in this collection 6

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 113 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

rl training for search agents squeezes exploration diversity while sft expands it — the same entropy collapse dynamic operates in search as in reasoning