Can abstractions guide exploration better than depth alone?
Does training a model to propose reasoning abstractions as intermediate subgoals help it explore diverse solution strategies more effectively than simply extending chain-of-thought depth?
RLAD addresses a structural problem with current reasoning training: RL incentivizes depth (longer chains attempting to verify one strategy) but not breadth (exploring diverse strategies). Long chains degenerate into frequent logic switches and unfocused exploration — the "underthinking" failure mode. Since Why do reasoning LLMs fail at deeper problem solving?, merely extending chains doesn't help.
The solution: reasoning abstractions — concise natural language descriptions of procedural and factual knowledge that function as high-level subgoals. Two models are jointly trained:
- Abstraction generator: given a problem, propose multiple reasoning abstractions (strategies, intermediate lemmas, relevant principles)
- Solution generator: conditioned on an abstraction, generate a solution that utilizes its information
The abstraction generator is rewarded for the improvement in solution accuracy that conditioning on its abstractions produces. The solution generator is rewarded for accuracy when using the abstraction. This cooperative two-player RL setup decouples learning signals: abstraction proposal and solution execution develop separately.
The key scaling result: allocating more test-time compute to generating abstractions is more beneficial for performance than generating more solutions — at large test budgets. This challenges the standard parallel sampling approach (generate N solutions, pick the best). Instead: generate diverse abstractions, then one good solution per abstraction. The abstractions enforce breadth where depth-only chains fail.
This connects to Why does parallel reasoning outperform single chain thinking? — abstractions are a mechanism for structured parallel exploration. And to Does separating planning from execution improve reasoning accuracy? — abstractions are a learned, RL-trained form of decomposition rather than a fixed prompt scaffold. In terms of the Can reasoning topologies be formally classified as graph types?, RLAD creates a two-level structure: parallel abstraction nodes (breadth-first, like CoT-SC) each conditioning a single depth-first solution chain (like CoT), producing a learned GoT-like topology where aggregation happens at the abstraction level.
The warmstart from SFT (summarize multiple candidate solutions → generate diverse abstractions) followed by RL refinement mirrors the Why does SFT-then-RL training follow a predictable three-phase pattern? dynamic, but in a cooperative multi-agent setting.
Inquiring lines that read this note 144
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Is language model reasoning authentic and what causes models to reason? Do language models develop actual world models or merely task heuristics?- Why do foundation models develop heuristics instead of world models?
- How do foundation models develop task-specific heuristics instead of world models?
- What distinguishes task-specific heuristics from genuine world models?
- How does iterative depth apply to world models and physical simulation?
- What scaffolding tools help users specify implicit contextual boundaries to models?
- Can graph cyclicity and topology predict when reasoning systems achieve breakthrough insights?
- How does nesting optimization levels improve on traditional network depth?
- What makes bilevel metacognition architectural rather than emergent in current systems?
- What role does exploration-exploitation balance play in abstraction formation?
- Does architectural design matter more than model scale for reasoning tasks?
- Can we transfer reasoning structure without copying surface form?
- Can recursive subtask trees implement tree-of-thought reasoning more efficiently?
- How does graph of thoughts enable divide-and-conquer reasoning patterns?
- Does small-world structure in reasoning graphs improve generalization?
- Can a single architecture represent both physical and mental possibility spaces?
- What makes a causal abstraction more transferable than a generic heuristic?
- How do progressive abstraction chains differ from branching reasoning topologies?
- Can weaker planners match stronger models if behavior is reorganized?
- Can surface heuristics override implicit constraints in domain-specific reasoning?
- Can marginal hints integrate better into reasoning than comprehensive explanations?
- Do depth thresholds correspond to transitions between procedural and strategic learning?
- What makes diverse reasoning sources more valuable than deeper single paths?
- Why do reasoning chains degenerate into undirected exploration at scale?
- Why do reasoning models wander instead of searching systematically?
- What distinguishes systematic search from wandering exploration in reasoning?
- Does verbal step-by-step reflection preserve learning signals that abstraction removes?
- Do reasoning failures stem from strategy or from calculation breakdown?
- Do reasoning models switch approaches when encountering local difficulty?
- What happens to iterative search quality when reasoning depth is unconstrained?
- How does Self-Discover compare to the cognitive tools approach?
- Why does the Chinese Room argument miss the deeper abstraction problem?
- When is numeric computation the real bottleneck versus reasoning depth?
- How does active reasoning through interaction differ from passive single-turn problem solving?
- How does o1-style reasoning relate to learned search processes versus memorized solutions?
- What makes multi-turn critique trajectories more effective than single-turn reasoning chains?
- Why does strategy diversity within reasoning chains improve model generalization?
- How does early commitment in reasoning differ from early exploitation in planning?
- Can explicit constraint statements override the dominance of surface heuristics?
- Why must procedural skills consolidate before strategic reasoning can develop?
- Do reasoning models trade instruction following for deliberative capability?
- Does reasoning structure match explicit versus implicit task demands?
- Why do models learn reasoning form instead of actual abstract inference?
- Is the reasoning cliff actually a tool-use problem?
- Do higher asymptote recipes unlock genuinely novel reasoning strategies?
- Why do foundation models develop task-specific heuristics instead of causal understanding?
- What is the relationship between reasoning depth and verbalization requirements?
- How much does chain-of-thought reasoning narrow the decompression gap?
- How does making implicit reasoning requirements explicit change model performance?
- How does SONAR embedding quality affect downstream reasoning accuracy?
- Why does explicit theory injection work better than example-based learning for reasoning tasks?
- Can activation patching reveal which reasoning steps actually matter?
- Does unrestricted reasoning per search step degrade iterative quality over time?
- Can reasoning improvements be attributed when optimizer and scaffold are unknown?
- Can step-level deliberation flags guide other reasoning systems?
- How do gradient descent iterations at inference compare to chain-of-thought reasoning chains?
- Why do longer reasoning chains signal hesitation rather than depth?
- Does deep-thinking ratio measure computational effort better than chain-of-thought length?
- How much reasoning depth do we actually need for most real-world tasks?
- Why does per-step deliberation lose global perspective compared to dynamic discovery?
- Do linearized traces genuinely expand exploration beyond standard chain-of-thought?
- Why do longer reasoning chains explore like tourists instead of scientists?
- What makes o1's chain-of-thought processing specifically effective for exploration tasks?
- Can tools unlock reasoning strategies that require abstract insight beyond computation?
- Does optimizing directly for semantic diversity improve both reasoning quality and exploration?
- Can prompting for specific creative paradigms improve ideation diversity?
- Can diverse critiques on a single problem unlock reasoning without diverse problem sets?
- How can semantic diversity optimization work if exploration and exploitation were truly opposed?
- Do novelty and feasibility always trade off in idea generation?
- Can a proposer agent actively surface a solver's weaknesses to prevent plateau?
- How would a bi-level agent restructure objective functions during discovery?
- How does semantic search over research papers guide autonomous architecture proposals?
- Do search agents face their own overthinking threshold like reasoning models do?
- Why do per-turn thinking budgets matter alongside iterative retrieval depth?
- How should AI ideation systems decompose and recombine research concepts?
- Why do per-turn reasoning caps improve iterative search quality?
- How does critique fine-tuning on one problem unlock broader reasoning?
- Why does imitation learning create a ceiling for reasoning capability?
- Why does extended reasoning training improve exploration without adding new capabilities?
- Does the model learn depth-wise drift as an explicit strategy?
- Does critique training improve exploration diversity during model training or only test time?
- How does policy initialization with sub-policies enable emergent thinking?
- What structural differences emerge between early generic skills and later meta-strategy skills?
- What makes exploration a verifiable and measurable training objective?
- How do evolutionary archives enable diverse exploration in self-improving systems?
- Can accelerated sampling techniques from image generation speed up evolutionary search?
- Can evolutionary approaches avoid the overthinking failure mode of iterative refinement?
- Can objective search escape the limitations of fixed-objective central planning?
- Can energy minimization replace reasoning-specific reinforcement learning for system 2 thinking?
- What is the optimal balance between search rounds and reasoning depth per round?
- What does an intermediate interface between planning and grounding actually look like?
- Does the planning-grounding factoring principle apply to other agent tasks?
- What makes planning, tool use, and reasoning into jointly optimizable subsystems?
- Can combinational creativity alone drive open-ended learning in agents?
- Can curriculum approaches teach agents when to stop exploring?
- Do task-specific heuristics emerge because they compress well enough?
- How does separating decomposition from execution improve multi-step reasoning?
- Can breadth-first search in continuous space outperform chain-of-thought on logical tasks?
- How does MCTS combine parallel exploration with sequential reasoning depth?
- When are multiple independent attempts more valuable than depth?
- Can width-scaling replace depth-scaling on inherently sequential problems?
- Can external summarization solve exploration problems in complex real-world environments?
- Do LLMs fail exploration because of context integration or computational limitations?
- How does explicit exploratory prompting compare to fine-tuned reinforcement learning for in-context adaptation?
- Why does prompting discover capabilities that need reward-driven refinement?
- How does inductive reasoning from partial evidence enable hypothesis formation?
- Why does exploration quality matter more than learner network depth?
- How does dynamic recurrence during training improve depth extrapolation?
- Can deterministic recurrent depth achieve the computational benefits of stochastic reasoning?
- What makes recursive depth more effective than parametric depth for puzzles?
- How does hierarchical recurrence compare to selective layer looping for computational depth?
- How does soft thinking compare to sampling multiple independent reasoning paths?
- How does continuous soft thinking explore multiple paths without explicit training?
- How do soft thinking and token-level mixtures explore multiple paths simultaneously?
- How does soft thinking achieve stochastic exploration without explicit training?
- Why is metacognition neglected as a foundational AI research area?
- How does the prefrontal cortex inspire artificial reasoning architectures?
- Does algorithmic decomposition prevent planning-execution interference in reasoning?
- How does planning-before-execution compare to iterative reasoning and action loops?
- Can backward planning reduce search difficulty when multiple goal state paths exist?
- How does separating decomposition from execution improve multi-step reasoning accuracy?
- How should humans specify deterministic abstractions of RL problems?
- Does RL amplify existing reasoning or create genuinely new computational strategies?
- Why does RL behavior differ between standard reasoning tasks and complex planning domains?
- Can the exploration ceiling be raised beyond what pretraining established?
- How do extrapolative and contextual generalization measure RL reasoning gains?
- How does interaction horizon differ from chain-of-thought depth?
- What distinguishes graph-of-thought reasoning from other structured reasoning topologies?
Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Why do reasoning LLMs fail at deeper problem solving?
Explores whether current reasoning models systematically search solution spaces or merely wander through them, and how this affects their ability to solve increasingly complex problems.
the problem RLAD addresses: depth without breadth
-
Why does parallel reasoning outperform single chain thinking?
Does dividing a fixed token budget across multiple independent reasoning paths beat spending it all on one long chain? This explores how breadth and diversity in reasoning compare to depth.
abstractions enforce structured parallel exploration
-
Does separating planning from execution improve reasoning accuracy?
Can modular LM architectures that split problem decomposition from solution execution outperform monolithic models? This explores whether decoupling these cognitive operations reduces interference and boosts performance.
abstractions as learned decomposition
-
Does policy entropy collapse limit reasoning performance in RL?
As reinforcement learning models become more confident in their policy choices, entropy drops and performance plateaus. Can we identify and counteract this bottleneck to sustain scaling?
abstractions may resist entropy collapse by maintaining strategy diversity
-
Can reasoning topologies be formally classified as graph types?
This explores whether Chain of Thought, Tree of Thought, and Graph of Thought represent distinct formal graph structures with different computational properties. Understanding this matters because the topology itself determines what reasoning strategies are possible.
RLAD creates a distinct topology: a two-level graph where the abstraction generator produces parallel breadth nodes (like CoT-SC) and each abstraction conditions a depth-first solution chain (like CoT); the result is a learned GoT-like structure where aggregation (in-degree > 1) happens at the abstraction level rather than at the solution level
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Reasoning LLMs are Wandering Solution Explorers
- RLAD: Training LLMs to Discover Abstractions for Solving Reasoning Problems
- ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
- Stream of Search (SoS): Learning to Search in Language
- Meta-Reasoner: Dynamic Guidance for Optimized Inference-time Reasoning in Large Language Models
- Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
- Algorithm of Thoughts: Enhancing Exploration of Ideas in Large Language Models
Original note title
reasoning abstractions decompose exploration into breadth-first strategy discovery and depth-first solution generation via two-player rl