Can evolutionary search beat sampling and revision at inference time?
Does population-based genetic search with LLM crossover and mutation outperform simpler inference strategies like best-of-N sampling and sequential refinement on natural language planning tasks?
Mind Evolution is an evolutionary search strategy for LLM inference that evolves a diverse population of candidate solutions. The LLM generates, recombines, and refines candidates based on evaluator feedback. This is analogous to combining divergent thinking (free-flowing parallel exploration) with convergent thinking (evaluation and selection) — considered hallmarks of intelligent problem-solving.
The key advantage over previous inference strategies: Mind Evolution works in natural language spaces without requiring task formalization. It only needs a programmatic solution evaluator — exploiting the observation that evaluating a candidate solution is often easier than generating one. This removes the need for formal problem definitions, expert-designed search spaces, or auxiliary verifiers.
Three mechanisms drive effectiveness:
- Population diversity via island model: Distinct sub-populations evolve independently between migration and reset events. Migration moves high-fitness solutions across islands; island reset replaces low-fitness populations with strong solutions from the global pool. This sustains exploration diversity that single-population evolution loses.
- LLM-based genetic operators: Instead of traditional mutation and crossover on symbolic representations, the LLM itself recombines and refines candidates using natural language understanding. This enables meaningful variation in unstructured solution spaces.
- Fitness-proportional selection: Parents with greater fitness are more likely to be selected for recombination, creating progressive quality improvement.
On TravelPlanner and Natural Plan benchmarks, Mind Evolution solves more than 98% of problem instances using Gemini 1.5 Pro — significantly outperforming Best-of-N and Sequential Revision when controlling for inference cost.
This extends the test-time compute landscape beyond the standard parallel-vs-sequential tradeoff. Mind Evolution is neither pure parallel sampling (Best-of-N) nor pure sequential refinement — it is iterative population evolution that combines elements of both. The island model specifically addresses the diversity collapse problem that Do iterative refinement methods suffer from overthinking? identifies — by maintaining multiple independent populations, evolution sustains exploration where single-trajectory refinement converges prematurely.
Inquiring lines that read this note 48
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Do language models develop actual world models or merely task heuristics? Do language models learn genuine understanding or just surface patterns? How do agent-learned skills transfer and improve across different tasks?- Do dynamic environments enable different kinds of agent-environment coevolution?
- Can RL-trained policies outperform text-space optimizers for evolving skill repositories?
- How do evolutionary archives enable diverse exploration in self-improving systems?
- Why do evolutionary algorithms collapse to single solutions under selection pressure?
- What makes diffusion sampling preserve multiple optimal solutions better than alternatives?
- How does latent space diffusion enable evolutionary search in high dimensions?
- Can accelerated sampling techniques from image generation speed up evolutionary search?
- Can evolutionary approaches avoid the overthinking failure mode of iterative refinement?
- How does fitness-proportional selection guide LLM recombination in unstructured solution spaces?
- Why does island model genetic evolution maintain diversity better than single populations?
- Does population-based evolution transcend the parallel versus sequential compute tradeoff?
- What distinguishes intrinsic search from extrinsic search method approaches?
- Is agentic efficiency analogous to convergent evolution in biology?
- Can evolutionary search unlock problems that best-of-n selection cannot solve?
- Can the same problem be solved by multiple evolutionary search strategies?
- Why does population-based search outperform both parallel and sequential test-time scaling?
- Why does test-time search also prioritize diversity over single-best convergence?
- Can objective search escape the limitations of fixed-objective central planning?
- Can LLM-based crossover and mutation work in unstructured natural language spaces?
- Does evolutionary inference transcend the parallel versus sequential test-time compute tradeoff?
- How does test-time search budget compare to evolution gains under matched conditions?
- Do evolutionary discovery systems like FunSearch count as bounded or open-ended improvement?
- What makes external diversity more effective than sequential revision steps?
- Can structural diversity through role assignment replace emergent diversity in small models?
- How does directional diversity compare to other forms of parallel planning?
- Why does genetic programming outperform direct LLM generation by 86 percent?
- Can optimization algorithms exploit the shift between procedural and planning bottlenecks?
- How should organizations redesign workflows if LLMs cannot solve optimization directly?
- How should we allocate model budget between evolvers and harness users?
- What feedback signals matter most during harness evolution search?
- Can harness evolution be redirected from memorization toward strategy distillation?
- How much of harness-evolution gain comes from matched test-time search budgets?
- Can harness evolution gains be distinguished from test-time search improvements on matched budgets?
Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Why does majority voting outperform more complex inference methods?
Simple majority voting across independent samples often matches or beats sophisticated alternatives like Best-of-N and sequential revision. What makes this basic approach so hard to beat for reasoning models?
Mind Evolution goes beyond voting: population-based recombination rather than just aggregation
-
Do iterative refinement methods suffer from overthinking?
Iterative refinement approaches like Self-Refine structurally resemble token-level overthinking in o1-like models. Does revision across multiple inference calls reproduce the same accuracy degradation seen within single inferences?
evolutionary approach avoids this through population diversity and island model
-
How should we balance parallel versus sequential compute at test time?
Test-time compute can prioritize breadth (trying many approaches) or depth (refining one approach). Which strategy works better, and does the answer depend on the problem?
Mind Evolution transcends this dichotomy: iterative evolution with parallel sub-populations
-
Can tree search replace human feedback in LLM training?
Explores whether Monte Carlo Tree Search can generate quality signals for self-improvement without expensive human annotations. Matters because annotation bottlenecks currently limit LLM scaling.
MCTS searches a tree; Mind Evolution searches a population; both use structured exploration
-
Does reward hacking always stem from the same failure?
Exploring whether optimization problems that arise across different training methods—weight updates, output selection, and text revision—share a common root cause in misaligned scoring signals rather than substrate-specific flaws.
selection by an evaluator over many generations is the frame's selection substrate; the frame ties exposure to how well search finds an evaluator's blind spots, and this note reports the evaluator as programmatic without saying whether it has any
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Evolving Deeper LLM Thinking
- Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution
- Self-Improving Language Models with Bidirectional Evolutionary Search
- A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?
- Learning to Discover at Test Time
- AgentGen: Enhancing Planning Abilities for Large Language Model based Agent via Environment and Task Generation
- Towards a Deeper Understanding of Reasoning Capabilities in Large Language Models
- Chain of Thoughtlessness? An Analysis of CoT in Planning
Original note title
evolutionary search at inference time outperforms best-of-n and sequential revision on natural language planning