SYNTHESIS NOTE
Topics›Deep Research›this note

Can models learn better by training on messy exploration paths?

Does including trial-and-error, reflection, and backtracking in training data teach models to reason more robustly than teaching only the polished shortest path to answers?

Synthesis note · 2026-06-03 · sourced from Deep Research
How does test-time scaling work for individual research agents?

Responding to OpenAI's opaque O1, this real-time replication effort contributes a paradigm beyond the engineering: journey learning. Where standard training teaches a model the shortcut — the clean path from problem to correct answer — journey learning encourages models to learn the complete exploration process: trial and error, reflection, and backtracking. The bet is that o1-style deep reasoning comes from internalizing how to search (including dead ends and recoveries), not from memorizing polished solution traces. The paper also models a methodological stance — transparent, continuously-documented, community-engaged research that reports failures as well as successes.

The keeper is the training-data philosophy: include the messy trajectory (failed attempts, self-correction) as the supervision signal, because that is what teaches robust reasoning, whereas shortcut-only data teaches confident-but-brittle answers.

This sits in the vault's reasoning-training thread. It is the constructive counter to the finding that Is reflection in reasoning models actually fixing mistakes? — journey learning tries to make exploration genuine rather than performative — and it pairs with When does RL actually extend reasoning beyond pretraining?: both concern what reasoning data actually teaches the model.

Inquiring lines that read this note 19

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How much do training data properties shape model reasoning? Why do stronger reasoning capabilities create tradeoffs with instruction following? What training dynamics and scale trigger emergence of reasoning capabilities? Why does adding new knowledge through fine-tuning degrade existing capabilities? What training data selection strategies maximize generalization across difficulty levels? What capability trade-offs arise from domain specialization through fine-tuning? How do soft reasoning mechanisms explore multiple paths without explicit training? Does RL create genuinely new reasoning capabilities or refine existing ones? Can inference-time compute effectively substitute for model scale? How do agent-learned skills transfer and improve across different tasks? Is reasoning capability latent in base models or created by post-training? What trajectory-level metrics beyond task success best evaluate agent performance? Can memory architectures handle ultra-long context better than attention? How does policy entropy collapse constrain scaling of reasoning-focused RL?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 157 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

journey learning trains models on the complete exploration process — trial error reflection and backtracking — not just shortcut solutions