Do reasoning traces need to be semantically correct?
Can models learn to solve problems from deliberately corrupted or irrelevant reasoning traces? This challenges assumptions about what makes intermediate tokens useful for learning.
"Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens" presents the strongest evidence yet against the assumption that reasoning traces carry meaningful semantics that contribute to solution quality.
The experimental design is clean. Transformers are trained on A* search traces for shortest-path planning in random mazes. Three conditions: (1) correct traces, (2) no traces, and (3) deliberately corrupted traces that have no relation to the specific problem they are paired with. The corrupted traces are not just noisy — they are systematically irrelevant, paired with wrong problems.
The results: corrupted-trace models maintain performance largely consistent with correct-trace models. In some cases they improve on correct-trace models and generalize more robustly to out-of-distribution tasks. Models trained on entirely correct traces still produce invalid reasoning traces when arriving at correct solutions — the formal A* validator confirms only a loose correlation between trace accuracy and solution accuracy.
This result directly challenges three assumptions simultaneously. First, that intermediate tokens function as reasoning steps (they may function as computational scaffolding — additional forward passes — regardless of semantic content). Second, that correct traces are superior training data (the scaffolding hypothesis predicts that any tokens providing additional computation would work). Third, that the "aha moment" in DeepSeek R1 indicates genuine realization (a single token insertion does not change internal state; it provides one more forward pass).
The "Stop Anthropomorphizing" position paper reinforces this from a different angle. It argues the community's tendency to call intermediate tokens "thoughts" or "reasoning traces" is actively harmful — generating false confidence and directing research toward improving trace quality rather than understanding the computational mechanism. The LLM-Modulo framework (generate-test with external verification) is proposed as the principled alternative: treat the LLM as a generator, use sound external verifiers for guarantees.
The practical implication: optimizing trace "interpretability" or "correctness" may be orthogonal to optimizing solution accuracy. The traces most useful for model performance may be those that provide optimal computational scaffolding, not those that most closely resemble human reasoning. This converges with What do models actually learn from chain-of-thought training?, which shows from the opposite direction that structural perturbations (shuffled steps) cause severe accuracy drops while content perturbations (wrong numbers, removed keywords) cause minimal impact. Together, these findings isolate the active ingredient: logical architecture, not semantic content.
Theoretical backing (RL-STaR): The theoretical analysis of the STaR framework provides formal support: RL-based self-taught reasoning can improve capabilities despite incorrect reasoning steps in the training data, because the iterative policy gradient converges under bounded error conditions. The model doesn't need correct intermediate steps to learn to produce correct final answers — what matters is the policy improvement trajectory, not the fidelity of individual traces. The quality of the pre-trained model sets the floor for effective bootstrapping, but the tolerance for noisy intermediates is built into the convergence guarantee.
Inquiring lines that read this note 310
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What causes reasoning models to fail or wander off track?- How can minimal pairs expose reasoning failures that single-instance accuracy metrics miss?
- What makes a background condition relevant to a specific reasoning task?
- Why do reasoning models fail on structurally unfamiliar instances?
- Can marginal hints integrate better into reasoning than comprehensive explanations?
- Why does explicit reasoning degrade passage reranking performance?
- What causes snowball errors to accumulate across reasoning steps in language models?
- How do failed branches remain in context and contaminate subsequent reasoning?
- Why does mixing reasoning traces from different teachers destabilize learning?
- Why does output alignment fail to catch internally incoherent reasoning?
- Can scaffolding frameworks isolate inductive reasoning from deductive confounds?
- Why do reasoning models wander instead of searching systematically?
- Why do larger reasoning models show cyclicity only in later layers?
- Why does cross-text analogical reasoning fail when semantics decouple from symbols?
- Why do reasoning models fail at learning hidden rules from sparse exceptions?
- What mechanisms cause reasoning models to wander rather than focus?
- How do single wrong steps corrupt entire reasoning chains?
- Can simple structure perturbations reliably expose memorization in reasoning models?
- Why do expert reasoners skip steps that novices must state explicitly?
- What failure modes emerge when scheme classification feeds downstream reasoning pipelines?
- Why might rationales that predict common text patterns fail on hard novel reasoning?
- Is reasoning failure caused by task complexity or training distribution gaps?
- How does recombining partial trajectories maintain coherence in natural language reasoning?
- Can outcome-focused objectives explain failures in reasoning evaluation?
- How does instance novelty rather than chain length explain reasoning failure?
- Why do structured reasoning representations sometimes reduce rather than improve error detection?
- Do recency-focused prompts and in-context examples work equally well for order recovery?
- Does irrelevant content degrade reasoning even when it fits the context window?
- Does irrelevant context degrade reasoning even within model context limits?
- Can corrupted reasoning traces be reliably distinguished from correct ones?
- Are correct reasoning traces measurably shorter than incorrect ones?
- How much accuracy is preserved when removing explanatory layers from reasoning traces?
- What behavioral markers signal when reasoning chains are performative?
- What makes a reasoning trace causally sufficient versus merely stylistically plausible?
- Why do correct reasoning traces appear shorter than incorrect ones?
- Can reasoning traces prove models are actually reasoning versus mimicking?
- How do planning and backtracking sentences control reasoning traces?
- Can concise reasoning traces match verbose explanation accuracy?
- Why do models show performative reasoning on easy tasks but genuine reasoning on hard ones?
- Can external verifiers replace reasoning trace quality in solution guarantees?
- Can removing failed branches from edited traces improve previous mistakes?
- Why are correct reasoning traces consistently shorter than incorrect ones?
- Can reasoning traces serve purposes beyond producing the final answer itself?
- Why do temporal reasoning patterns matter more than final answers?
- Does logical trace coherence guarantee valid mathematical reasoning?
- Does reasoning trace style explain why RL post-training improves model reasoning?
- Do shorter reasoning traces actually produce more reliable model outputs?
- Can synthesized explanations be more auditable than winning-chain explanations?
- Why do correct reasoning traces tend to be shorter than incorrect ones?
- Why does intermediate step quality predict reasoning outcomes better than global features?
- Why do corrupted traces maintain performance as well as correct traces?
- How does post-training on traces improve performance without semantic reasoning?
- Does anonymizing reasoning traces harm the quality of model outputs?
- Which sentences in reasoning traces actually influence the final answer?
- Why do invalid reasoning steps produce nearly the same performance gains?
- Can deliberate corruption of reasoning traces harm out of distribution generalization?
- Why do reasoning models produce unfaithful or unhelpful reasoning traces?
- Why do invalid prompts produce reasoning traces as effectively as valid ones?
- Why do reasoning traces resemble mimicry rather than verified problem-solving?
- What distinguishes coherent reasoning from inaccurate but plausible predictions?
- How does trace coherence differ from valid mathematical proof in practice?
- How does trace coherence differ from trace validity in reasoning?
- Can models maintain auditable reasoning while achieving high accuracy?
- Do correct reasoning traces tend to be shorter than incorrect ones?
- What makes some sentences in reasoning traces have disproportionate causal influence?
- Why do models skip steps that would make reasoning clearer?
- Do shorter correct reasoning traces contain more thought anchors than longer ones?
- Do corrupted reasoning traces teach something different than pure success traces?
- Why does failed step fraction predict reasoning quality better than trace length?
- Why do correct reasoning traces stay shorter than incorrect ones?
- Why are incorrect reasoning traces longer than correct ones?
- What role do local backtracking steps play in reasoning traces?
- Does trace length actually reflect problem difficulty or training proximity?
- Why do wrong numbers cost less accuracy than shuffled reasoning steps?
- Do longer chain-of-thought traces improve interpretability or just performance?
- Why do reasoning traces mislead users into trusting wrong model answers?
- How much of a reasoning trace is actually redundant or unnecessary?
- How can reasoning quality be verified before integrating new information into a reasoning graph?
- What distinguishes genuine capability gains from coherent but invalid reasoning traces?
- Why do reasoning traces persuade users without improving their accuracy?
- What quality filters distinguish useful reasoning enrichment from shallow repetition?
- How much do compressed reasoning traces transfer across different models?
- What makes a thinking trace take information shortcuts?
- Why do shorter confident reasoning traces fail on out-of-distribution problems?
- Can post-hoc analysis of reasoning traces actively mislead users?
- Why do language model reasoning chains look fluent when they deviate from the task?
- What makes reasoning traces effective or ineffective for solving problems?
- Why do corrupted reasoning traces sometimes generalize better than correct ones?
- How does confidence filtering improve selection of reasoning traces?
- Why are shorter reasoning traces more reliable than longer correct ones?
- What makes some reasoning traces better supervision than others despite equal accuracy?
- Why do reasoning traces fail to accurately reflect model decision-making?
- Why do language models produce reasoning traces that mimic human reasoning style?
- Is the structure of reasoning traces learned as a shared stylistic convention?
- Can reasoning traces reliably distinguish genuine value conflicts from reasoning errors?
- How do you supervise reasoning that never becomes tokens?
- Why do invalid reasoning prompts work as well as valid ones?
- Can flow concentration in reasoning traces predict model quality better than tokens?
- Why do deliberately corrupted reasoning traces sometimes generalize better than correct ones?
- What makes a set of traces collectively useful beyond their individual quality?
- Why do reasoning models produce unfaithful derivational traces by default?
- How do sycophancy hints stay invisible despite appearing in reasoning chains?
- How do reasoning traces serve as hypotheses about decision processes?
- Can problem structure and representation format be mismatched intentionally?
- Can reasoning traces that feel convincing fail to help people predict behavior?
- Does faithfulness in reasoning traces guarantee people can verify model outputs?
- Can reasoning traces be verified for authentic single authorship?
- Why does scaling reasoning tokens fail to improve unfamiliar tasks?
- Do tokens beyond a critical threshold actually improve reasoning quality?
- Does thinking-token overuse actually degrade reasoning accuracy in practice?
- Why does reasoning accuracy degrade beyond a critical thinking token threshold?
- Can token efficiency come from stopping before reflection?
- How much of a model's reasoning tokens are unnecessary for reaching the final answer?
- How does constraint complexity relate to optimal reasoning token budgets?
- What happens to reasoning accuracy when models use more thinking tokens?
- Why do reasoning models reduce effort despite having token budget remaining?
- How much does test-time compute improve reasoning without more tokens?
- Can early stopping on reflection tokens save computation without accuracy loss?
- Why does more inference compute amplify wandering rather than solving it?
- Why does representation recycling of MI-peak tokens improve reasoning accuracy?
- What happens to model reasoning accuracy as thinking token requirements exceed critical thresholds?
- Why does uniform averaging across all tokens dilute the reasoning signal?
- What causes reasoning accuracy to degrade beyond a critical thinking-token threshold?
- How does SONAR embedding quality affect downstream reasoning accuracy?
- Why does explicit theory injection work better than example-based learning for reasoning tasks?
- How does optimizing for accuracy during training degrade downstream reasoning quality?
- Can verifier-guided search catch factual errors that reasoning training cannot?
- Why do open-source models trained on proprietary outputs still fail at reasoning?
- Why does domain accuracy improve while reasoning quality degrades after supervised fine-tuning?
- Does supervised fine-tuning improve accuracy while damaging the quality of reasoning?
- How can entailment benchmarks separate genuine reasoning from memorization effects?
- Do reasoning models perform genuine logical evaluation or pattern matching?
- Why does supervised fine-tuning degrade reasoning quality despite raising accuracy?
- Why do SFT models memorize patterns instead of learning generalizable reasoning?
- How does data quality mismatch create reasoning degradation in supervised fine-tuning?
- Can reasoning catalyst data serve as a stable foundation for test-time training?
- Why does eliminating proxy-model filtering improve reasoning emergence in pretraining?
- Why do reasoning tasks improve more than retrieval from lookup memory?
- Why does naive randomness fail to improve stochastic latent reasoning models?
- How does contrapositive augmentation change the tractability of reasoning tasks?
- Does reasoning efficiency transfer to tasks without ground truth dependency graphs?
- Can reasoning improvements be attributed when optimizer and scaffold are unknown?
- Why do single examples trigger large reasoning improvements in models?
- Can models learn to select exemplars based on reasoning skills rather than complexity?
- How does business logic specification replace annotated training datasets?
- How much reasoning catalyst data is actually needed for improvement?
- How does factoring perception from reasoning improve sparse-label learning?
- Why do recursive belief models require different training than logical derivation?
- Can training improve reasoning coherence without improving actual correctness?
- How does a single training example trigger phase transitions in reasoning output?
- Can a single correct example seed exponential improvement in mathematical reasoning?
- How can one training example improve reasoning across thousands of unseen problems?
- Does latent reasoning capability exist in base models before any training?
- How does backward reasoning during training improve forward reasoning capability?
- Can models reason at inference without specialized internal training?
- Why does reasoning transfer across different numbers but factual recall does not?
- Can smaller amounts of diverse reasoning demonstrations replace exhaustive factual training data?
- What makes token-level reasoning during pretraining different from test-time chain-of-thought?
- Does token-level reasoning during pretraining improve general reasoning without task-specific supervision?
- Does the base model already contain latent reasoning capability?
- Can models possess latent reasoning capability that training signals fail to unlock?
- Why do knowledge and reasoning train in different network layers?
- Can approximate or noisy reference answers work for RL-based reasoning training?
- Why does reasoning backward enable better forward reasoning performance?
- Can minimal training signals unlock latent reasoning capability in base models?
- Can minimal training signals unlock reasoning already latent in pretrained representations?
- What latent reasoning capability do base models already possess before training?
- Can reflection in reasoning models be corrective rather than just confirmatory?
- Do self-revision tokens measurably degrade reasoning accuracy in scaled models?
- Why does iterative refinement amplify rather than correct reasoning errors?
- Why does reflection in reasoning models stay confirmatory instead of corrective?
- Can training on reasoning traces teach actual self-correction or only confident first answers?
- Why does reflection in reasoning models tend to be confirmatory rather than corrective?
- What inference strategy works better than forcing self-revision under token constraints?
- Why does reflection in reasoning models confirm rather than correct initial directions?
- How do prior errors in reasoning context amplify future mistakes?
- Do reasoning models need to verbalize doubt to correct their own mistakes?
- Can latent reasoning architectures work as retrofits to existing models?
- Can latent reasoning mechanisms and recursive tracking mechanisms be combined effectively?
- Can latent reasoning achieve the same substitution without tokens?
- Can articulating latent reasoning processes improve transfer across domains?
- Why does latent-level prediction beat token-level prediction for reasoning?
- Can latent reasoning scale test-time compute without verbal tokens?
- Can latent reasoning scale test-time compute without verbalized tokens or special training?
- Do models leak their true associations through reasoning traces and behavior?
- Does the latent-explicit gap widen beyond 3B parameters on reasoning tasks?
- Can latent reasoning stay readable without explicit token-by-token decoding?
- Can models learn when to invoke search during reasoning tasks?
- Why do models learn reasoning form instead of actual abstract inference?
- Can models distinguish between activated knowledge and genuine reasoning?
- Is the reasoning cliff actually a tool-use problem?
- Why do difficult problems force models to develop reasoning strategies?
- Can reasoning models succeed at logic but fail at execution?
- Can models be trained to explain instead of imitate answers?
- Why do familiar patterns that support correct answers sometimes drive errors?
- Why do reasoning model failures stem from execution rather than reasoning?
- Can training models on backward reasoning improve their forward planning ability?
- Why do students learn better from explanations than from solving problems from scratch?
- How can we turn reasoning model failures into useful training signals?
- Can instruction-level interventions fix memory-induced reasoning failures in practice?
- Does iterative denoising order affect the reasoning style diffusion models learn?
- Can we measure how much prior errors bias subsequent token predictions?
- How does token-by-token generation constrain a model's ability to plan ahead?
- Do reflection tokens and symbolic tokens serve different roles in reasoning?
- How early in token generation does the reasoning mode activate?
- How does tokenization change what gets counted as valuable knowledge?
- What distinguishes memorized tokens from causally necessary reasoning steps?
- Can knowledge density per token be measured as a quality metric?
- Which tokens actually change across different reasoning paths in rollouts?
- How do reasoning-invariant tokens dilute learning signals in uniform averaging?
- Can models internally identify which tokens matter most for reasoning?
- How do meta-tokens help models learn when to generate reasoning versus commit predictions?
- How does predictive accuracy on future tokens differ from correctness on labeled answers?
- What evidence shows that reasoning chains encode token-level functional structure?
- Does the token prediction framing actually capture what human reasoning does?
- Can this whole-artifact principle apply to other generative tasks?
- Why do language models use remaining tokens to rationalize instead of reconsider?
- Are reasoning traces really reasoning or just stylistic imitation of human thought?
- Why do logically invalid chain-of-thought examples work nearly as well?
- Can chain of thought traces be designed to prevent anthropomorphic misinterpretation?
- Why does chain-of-thought fail when problems lack matching training schemata?
- Does the DeepSeek R1 single token insertion represent genuine reasoning?
- How do explicit reasoning traces help models construct valid syntactic trees?
- Why do we measure reasoning quality by reading visible chains?
- Why do verbalized reasoning chains fail on certain problem classes?
- Can chain-of-thought traces be faithful without causal sufficiency and necessity?
- Why do models rarely admit to their actual reasoning in chain-of-thought traces?
- Can instance-adaptive reasoning happen without sequential token dependencies?
- Are chain-of-thought traces anthropomorphizing how AI models really reason?
- Can chain-of-thought traces harm rather than help user understanding?
- Does reasoning training create blind spots in premise detection?
- Why does chain-of-thought monitoring fail on mixed-authorship reasoning traces?
- Why do correct reasoning traces in language models tend to be shorter?
- When does explicit reasoning actually degrade performance on a task?
- Why does extending reasoning traces worsen persona consistency?
- Can inserted errors in reasoning drafts produce predictable downstream effects?
- Can memorization scores diagnose where reasoning chains become unreliable?
- Do linearized traces genuinely expand exploration beyond standard chain-of-thought?
- Can we detect redundant reasoning steps during model inference instead of training?
- How much does training data format shape what reasoning strategy emerges?
- How does training data format shape which reasoning patterns emerge in models?
- Why does training data format shape reasoning strategy more than content?
- Can training format itself shape what reasoning strategy a model learns?
- Does partial trace guidance work better than curriculum learning for hard problems?
- Can models learn better from critiquing errors than imitating correct responses?
- Can partial solution traces convert unproductive hard samples into learnable training data?
- Can solution traces substitute for process-level reward signals in math reasoning?
- Why does outcome supervision fail for long reasoning chains?
- Can small models solve complex tasks using externalized reasoning graphs?
- Why does reasoning graph topology evolve differently across training phases?
- What makes training data quality more important than quantity for reasoning?
- What makes some training data teach brittle answers versus robust reasoning?
- Why do entities trigger memorized propositions instead of enabling reasoning?
- Can derivational traces be distinguished from stylistic mimicry of reasoning?
- Can LLM reasoning traces be validated against actual population reasoning?
- Why is extracting training data insufficient proof that models memorize?
- Why does grokking reveal the shift from memorization to genuine understanding?
- Why does fine-tuning models for continuous reasoning cause catastrophic forgetting?
- Can models recover knowledge with completely unrelated retraining tasks?
- Can event boundaries be identified from statistical regularities without understanding events?
- Can mechanistic interpretability explain explanation-execution disconnection?
- Why does distillation transfer reasoning patterns with few examples?
- Does reasoning style transfer matter more than solution correctness in distillation?
- What training signals would teach models when not to reason?
- Why do reasoning models confidently generate wrong answers instead of abstaining?
- Can reasoning models reject ill-posed questions or do they overthink?
- Why do language models generate reasoning tokens after internally deciding the answer?
- Why do explicit linguistic markers override semantic computation in models?
- What sparse mechanistic structures drive reasoning traces in language models?
- What distinguishes inductive inference from negative evidence versus positive patterns?
- How does correctness emergence occur when no expert initially solved the task?
- Why does the order of training examples matter for what models learn?
- What specific tasks should evaluate whether models understand pedagogical sequencing?
- How can correct explanations coexist with failed applications in AI?
- How should tool-call attribution distinguish credit between successful accidents and intentional actions?
- How do continuous concept tokens compare to latent trajectory sampling?
- How do continuous concept tokens explore multiple reasoning paths without explicit sampling?
- Can base models spontaneously produce reasoning traces without any RL training?
- Why does standard RL cause traces to collapse into redundant reasoning paths?
- Do reasoning traces actually make better reward models for grading answers?
- How does reward density during training affect token efficiency in reasoning?
Related concepts in this collection 8
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Do reasoning traces actually cause correct answers?
Explores whether the intermediate 'thinking' tokens in R1-style models genuinely drive reasoning or merely mimic its appearance. Matters because false confidence in invalid traces could mask errors.
this provides the strongest evidence for the stylistic mimicry claim: even irrelevant mimicry works
-
Does chain-of-thought reasoning reveal genuine inference or pattern matching?
Explores whether CoT instructions unlock real reasoning capabilities or simply constrain models to mimic familiar reasoning patterns from training data. This matters for understanding whether language models can actually reason abstractly.
extends: not just constrained imitation but imitation of form without semantic content still effective
-
Do language models actually use their reasoning steps?
Chain-of-thought reasoning looks valid on the surface, but does each step genuinely influence the model's final answer, or are the reasoning chains decorative? This matters for trusting AI explanations.
extends necessity failure: traces can be semantically irrelevant and still produce correct solutions
-
Can minimal reasoning chains match full explanations?
Does removing all explanatory text from chain-of-thought reasoning preserve accuracy? This tests whether verbose intermediate steps are necessary for solving problems or just artifacts of how language models are trained.
both findings converge: most trace content is dispensable
-
Does training on messy search processes improve reasoning?
Can language models learn better problem-solving by observing full exploration trajectories—including mistakes and backtracking—rather than only optimal solutions? This matters because current LMs rarely see the decision-making process itself.
complementary finding: corrupted traces show content is dispensable (scaffolding hypothesis); SoS shows the search PROCESS itself is valuable training data (process exposure hypothesis). Different mechanisms, both challenge optimal-trace supremacy.
-
What do models actually learn from chain-of-thought training?
When models train on reasoning demonstrations, do they memorize content details or absorb reasoning structure? Testing with corrupted data reveals which aspects of CoT samples actually drive learning.
convergent evidence from the opposite direction: corrupted content is tolerated (this note) while corrupted structure causes severe degradation (that note); together they confirm traces function as structural scaffolding not semantic reasoning
-
Does logical validity actually drive chain-of-thought gains?
What if invalid reasoning in CoT exemplars still improves performance? Testing whether logical correctness or structural format is the real driver of CoT's effectiveness.
convergent finding from prompting rather than training: invalid exemplar reasoning at inference time (that note) parallels corrupted training traces (this note), both showing logical validity is dispensable for performance gains
-
What three separate factors drive chain-of-thought performance?
Can we isolate and measure the distinct contributions of output probability, memorization, and genuine reasoning to CoT success? Understanding their relative weights matters for knowing when CoT actually reasons versus when it relies on shortcuts.
the three-factor decomposition explains WHY corrupted traces work: output probability (the dominant factor) is shifted by intermediate token generation regardless of content validity; only the noisy-reasoning factor requires semantic correctness
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens
- Performative Thinking? The Brittle Correlation Between CoT Length and Problem Complexity
- Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
- Local Coherence or Global Validity? Investigating RLVR Traces in Math Domains
- Faith and Fate: Limits of Transformers on Compositionality
- Reasoning Can Hurt the Inductive Abilities of Large Language Models
- Beyond the Last Answer: Your Reasoning Trace Uncovers More than You Think
Original note title
deliberately corrupted reasoning traces perform comparably to correct traces and sometimes generalize better out of distribution