Does logical validity actually drive chain-of-thought gains?
What if invalid reasoning in CoT exemplars still improves performance? Testing whether logical correctness or structural format is the real driver of CoT's effectiveness.
"Invalid Logic, Equivalent Gains" runs a clean experiment: replace valid reasoning in CoT exemplar prompts with completely illogical reasoning, then measure performance on BIG-Bench Hard tasks. The result: logically invalid CoT prompts perform close behind valid CoT and outperform answer-only prompting. The reasoning content of CoT exemplars is not what drives the performance gain.
This is a sharp test because it isolates the contribution of logical validity from everything else CoT provides: output format, step decomposition, intermediate token generation, attention pattern scaffolding. If invalid reasoning still helps, then the benefit comes from these structural properties, not from the reasoning itself.
The finding directly supports Does chain-of-thought reasoning reveal genuine inference or pattern matching?. If the model were learning to reason from exemplars, invalid exemplars would degrade performance substantially. Instead, the model is learning the FORM of step-by-step output — the structure activates latent capabilities without the exemplar content needing to be logically sound.
This also deepens Do language models actually use their reasoning steps?. If the exemplar reasoning doesn't need to be valid for CoT to work, then the model's own generated reasoning may similarly be decorative rather than causal. The exemplar finding makes the faithfulness concern bidirectional: neither the input reasoning (exemplars) nor the output reasoning (generated CoT) need be logically valid for the performance gain to occur.
The practical implication: CoT prompt engineering should focus on structural properties (step count, decomposition format, answer scaffolding) rather than on the logical correctness of the exemplar reasoning. Since Why do chain-of-thought examples fail across different conditions?, the dimensions that matter are structural (complexity, order, style), not logical.
Inquiring lines that read this note 231
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do false presuppositions and sycophancy drive persistent false beliefs in models?- What makes counterfeiting social warrant different from counterfeiting factual claims?
- What makes a claim socially valid even if factually imprecise?
- Why do humans trust explanations that fail counterfactual prediction tests?
- Can formal argumentation structure replace ad-hoc fallacy classifications?
- How do belief edits differ between surface endorsement and deep integration?
- Can a single fabricated evidence payload shift model beliefs without multi-turn pressure?
- How does validation skill replace production skill in AI systems?
- What structural features force users to evaluate the epistemic status of outputs?
- What structural evidence shows that polished presentation substitutes for actual thinking in AI output?
- How do satisfaction scores differ from genuine cognitive improvement?
- Do explicit reasoning formats help or hurt human judgment across tasks?
- Why do people accept generated output that sounds convincing but lacks support?
- What makes accountability and validity-orientation non-behavioral properties?
- How does cognitive load explain linguistic patterns in both deception and incorrect reasoning?
- Can users reliably distinguish valid reasoning from plausible-looking deception?
- What makes emotional alignment more effective than logic when reasoning errors are exposed?
- What makes quasi-beliefs real enough to explain AI behavior?
- How does externalizing tacit expertise into structured rules differ from prompt engineering?
- Does good simulation eventually count as genuine realization?
- What are the seven components of genuine mental state simulation?
- Can functional behavior alone capture what makes something a genuine belief?
- Can a perfect behavioral simulation constitute genuine understanding or experience?
- Why does item discrimination matter more than surface-level question plausibility?
- Can contextual design decisions resist formalization into evaluation rubrics?
- How does evaluation format change what we measure about model reasoning?
- Do current math benchmarks measure outcomes or rhetorical plausibility?
- Why does sophisticated measurement not validate the underlying scientific inference?
- Can proxy evaluation of ideas accurately predict their quality without implementation?
- Does meta-judging improve evaluator quality better than temporal decoupling alone?
- What makes Beck's diagram effective for constraining simulated patient behavior?
- What makes clinical theory grounding more effective than pattern matching alone?
- How much does faithfulness vary naturally in reasoning without evaluation pressure?
- What detection methods can catch each distinct CoT bypass strategy?
- Why do logically invalid chain-of-thought examples work nearly as well?
- Does each reasoning step in chain-of-thought introduce cumulative error?
- What happens to chain-of-thought performance across distribution shifts?
- Can chain-of-thought reasoning be genuinely causal if exemplars don't need logic?
- Does chain-of-thought reasoning amplify bullshit or just make it more visible?
- How do we verify that stated beliefs actually follow from underlying motifs?
- What three factors actually drive chain of thought performance improvements?
- Why do we measure reasoning quality by reading visible chains?
- Why do chain-of-thought outputs look logical but perform rhetorically?
- Why does long CoT training optimize for structural coherence over content correctness?
- How does chain of thought amplify specific forms of rhetorical bullshit?
- Does optimizing against CoT monitors inevitably produce obfuscated reasoning?
- How does faithfulness differ from informativeness in chain-of-thought evaluation?
- Can chain of thought monitoring reliably catch model misbehavior?
- What distinguishes metacognitive regulation from standard chain-of-thought reasoning?
- Does CoT reasoning actually cause the outputs that follow it?
- How do thought actions represent policy improvement steps in practice?
- How brittle are chain-of-thought exemplars across order and complexity?
- Can chain-of-thought disclosure measure whether reviewers actually notice model errors?
- Do recency-focused prompts and in-context examples work equally well for order recovery?
- Does irrelevant content degrade reasoning even when it fits the context window?
- How do output format constraints compare to input exemplar brittleness?
- Can structured prompts reduce reasoning steps while improving financial accuracy?
- Can operationalizing theory into prompt structure improve reasoning more than theory itself?
- What distinguishes planning knowledge from an executable plan that works?
- Can structured decomposition fix evaluation gaps in other research tasks?
- What makes training-free approaches like Soft Thinking preferable to SoftCoT?
- Can models learn to select exemplars based on reasoning skills rather than complexity?
- Can reasoning skills trained on law improve performance in STEM?
- Can training improve reasoning coherence without improving actual correctness?
- Can a single correct example seed exponential improvement in mathematical reasoning?
- Can thought quality alone be trusted to guide model training?
- How does RPT compare to learning when versus how to deploy reasoning?
- What makes the verifier the load-bearing component of reasoning training?
- What would whole-system AGI evaluation look like in practice?
- How do surface correlations between narratives and answers mislead benchmark validity?
- When does the correlation between consistency and correctness break down?
- Does the verification gap widen exactly where judgment replaces checkability?
- How should process quality and verification cost factor into evaluation judgment?
- How do live human evaluations differ from ground-truth benchmarks?
- What structural changes help AI generation keep pace with verification?
- Why do contrastive reasoning approaches outperform single-path belief evaluation?
- What makes Compound-QA expose weaknesses in monologue reasoning?
- What explains the gap between perplexity performance and actual reasoning capability?
- Can simple structure perturbations reliably expose memorization in reasoning models?
- Why does the Chinese Room argument miss the deeper abstraction problem?
- Can outcome-focused objectives explain failures in reasoning evaluation?
- What makes schema identification necessary after assessing thoughts and evidence?
- Why might expressed satisfaction with explanations diverge from actual cognitive clarity?
- Can the three-stage DoT framework detect all cognitive distortion types reliably?
- Why do user studies of explanations fail to predict deployed effectiveness?
- How does cognitive fit theory explain why different tasks need different knowledge structures?
- How much does training composition affect syntactic versus reasoning performance?
- How large was the effect size of format compared to content itself?
- How much do mechanistic interpretability findings reflect true reasoning architecture?
- How do mechanistic interpretability tools help distinguish truthfulness from honesty?
- Can high-entropy tokens and step-level confidence identify the same critical reasoning forks?
- Do gold CoT tokens avoid the need for specialized training data?
- Why do benchmark designers treat content effects as confounds?
- Can high test performance mask a complete absence of understanding?
- What evaluation methods actually measure reasoning versus execution capability?
- How can benchmark accuracy scores mask the absence of interpretable reasoning structure?
- How much of MATH-500 improvement comes from data contamination versus real reasoning gains?
- Can memorization inflate apparent capability on benchmarks with available solutions?
- How much of weak-to-strong performance gaps reflect presentation rather than capability?
- What mechanism causes confident false answers under high cognitive load?
- What makes accurate confidence different from confident-but-wrong predictions?
- How does model confidence relate to exemplar brittleness in chain-of-thought?
- What makes mathematically confident but incorrect answers resemble valid solution shapes?
- How do confident system outputs weaken user skepticism about their reliability?
- Why is faithful calibration considered fundamentally metacognitive?
- Can reasoning benchmarks separate logic from believability?
- Can activation patching reveal which reasoning steps actually matter?
- How can entailment benchmarks separate genuine reasoning from memorization effects?
- Do reasoning models perform genuine logical evaluation or pattern matching?
- Does SFT degrade reasoning quality while improving domain accuracy?
- Can reasoning evaluation metrics reward actual reasoning instead of theater?
- Can reasoning catalyst data serve as a stable foundation for test-time training?
- Why do benchmark scores rise while reasoning quality declines?
- Why does contextual judgment matter more in law and medicine than in mathematics?
- Do explicit reasoning chains improve or harm performance on complex judgment tasks?
- Why do top performers produce shorter chains of thought in their strongest domains?
- Does explicit reasoning help or hurt tasks requiring continuous nuanced judgment?
- How does difficulty level change whether extended thinking provides genuine reasoning signal?
- When does explicit reasoning actually degrade performance on a task?
- How do chain-of-thought structures affect reasoning robustness?
- Why does step-by-step reasoning degrade performance on judgment-based tasks?
- Why does extended thinking increase output variance without improving reasoning quality?
- Does explicit reasoning help or hurt tasks requiring continuous judgment?
- Does the answer stage perform substantial reasoning beyond the thinking draft?
- Does performative reasoning mask underlying uncertainty even on easy problems?
- How much does training data format shape what reasoning strategy emerges?
- How does training format shape reasoning strategy more than content?
- How much does training data presentation format shape reasoning ability?
- Why does training data format shape reasoning strategy more than content?
- Can reasoning chains work without logical validity?
- What makes symbolic operations different from general knowledge questions?
- What makes structural logic correlate so strongly with contextual consistency?
- Why do format and structure matter more than actual content in reasoning?
- Why does premise ordering shift syllogistic reasoning performance by over 30 percent?
- Why does augmenting symbolic reasoning outperform replacing it entirely?
- Can structured reasoning replace execution for runtime behavior verification?
- What makes structured informal reasoning preferable to full formalization?
- Can test-time scaling prioritize genuine reasoning over pattern matching?
- What patterns emerge across test-time scaling and reasoning architectures?
- What makes counterfactual thinking different from behavioral pattern matching?
- What are collider structures and why do they reveal reasoning errors?
- How does vehicle causality differ from content causality in physical systems?
- What saliency patterns distinguish successful from failed chain-of-thought reasoning?
- Does logical trace coherence guarantee valid mathematical reasoning?
- Can synthesized explanations be more auditable than winning-chain explanations?
- What attention mechanisms explain why verification steps get ignored?
- Why do invalid reasoning steps produce nearly the same performance gains?
- Why do invalid prompts produce reasoning traces as effectively as valid ones?
- What distinguishes coherent reasoning from inaccurate but plausible predictions?
- What makes well-formatted outputs misleading as evidence of model capability?
- Why do semi-formal templates improve verification accuracy over unstructured reasoning?
- How does test-time verification decouple the act of checking from reasoning generation?
- What distinguishes genuine capability gains from coherent but invalid reasoning traces?
- Why do invalid reasoning prompts work as well as valid ones?
- What role does verifier design play in reasoning capability gains?
- Can problem structure and representation format be mismatched intentionally?
- Can reasoning traces that feel convincing fail to help people predict behavior?
- Can structured output formats reduce instruction following degradation?
- Do synthetic verification chains from long-CoT models match the quality of human-annotated process labels?
- Can correct outputs mask reliance on surface heuristics rather than deep understanding?
- Why do instruction following and reasoning capability trade off in training?
- Can you steer reasoning by directly manipulating SAE features?
- Can format adaptation alone explain why reasoning enrichment improves instruction following?
- Why does target probability matter more than task logical complexity?
- How should researchers evaluate whether correct model outputs reflect real structural learning?
- Why does the gap between theoretical expressiveness and learned capability matter?
- What distinguishes inductive inference from negative evidence versus positive patterns?
- How does training on correct answer form differ mechanistically from training on failure analysis?
- What structural differences emerge between early generic skills and later meta-strategy skills?
- Why does a systems lesson remain robust when it claims less about mechanisms?
- What qualities make a behavioral pattern count as a teachable skill?
- Do fed-back concepts or the auxiliary objective alone drive the performance gain?
- Can chain of thought reasoning actually validate logical arguments?
- What types of math proofs benefit most from proof-by-contradiction framing?
- Which structural properties of CoT prompts matter most for performance?
- How do exemplar properties affect the brittleness of chain-of-thought prompting?
- Why does reasoning effort fail to improve theory of mind performance?
- How do structured benchmarks hide theory of mind failures in LLMs?
- Why does additional reasoning effort not improve theory of mind performance?
- Does reasoning effort correlate with social reasoning accuracy?
- Can structured theory of mind benchmarks measure genuine mental state reasoning?
- How can we measure whether process rewards actually align with reasoning quality?
- How do partial credit grading systems accidentally reward reasoning theater?
- Can multiple verification approaches together overcome the self-improvement ceiling?
- What distinguishes intrinsic metacognition from extrinsic human-designed loops?
- Can held-out validation gates prevent optimizer hallucinations in skill proposals?
- How do test harnesses guide reflection better than transcripts alone?
- Does belief-shift credit assignment generalize to tasks without ground-truth outcomes?
- How can structured reasoning templates serve as rewards for code agent training?
- What role does task structure play in rewarding delayed thinking?
- Why do explicit quality criteria outperform learning quality from examples alone?
- Why does exemplar performance vary across order complexity diversity and style?
- Can skill validation through testing prevent unreliable programs from accumulating?
- Does inspectable skill artifacts guarantee the behavior matches the person it claims to ground?
- What makes a deployment paradigm credible for maintaining scientific integrity?
- Can a correct outcome hide a fundamentally unsound decision-making process?
- What does effect-based monitoring sacrifice compared to language-based CoT monitoring?
- Why does showing counterarguments restore users' ability to discriminate?
- Why does inference-time debate fail when persuasion substitutes for evidence?
- Does sounding confident in framing make arguments more persuasive despite weaker logic?
- Do form, evidentiality, and tone interact with the size effect?
Related concepts in this collection 6
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does chain-of-thought reasoning reveal genuine inference or pattern matching?
Explores whether CoT instructions unlock real reasoning capabilities or simply constrain models to mimic familiar reasoning patterns from training data. This matters for understanding whether language models can actually reason abstractly.
invalid exemplars still working confirms form-over-content thesis
-
Do language models actually use their reasoning steps?
Chain-of-thought reasoning looks valid on the surface, but does each step genuinely influence the model's final answer, or are the reasoning chains decorative? This matters for trusting AI explanations.
bidirectional unfaithfulness: exemplar validity and output validity both decorative
-
Why do chain-of-thought examples fail across different conditions?
Chain-of-thought exemplars show surprising sensitivity to order, complexity level, diversity, and annotator style. Understanding these brittleness dimensions could reveal what makes reasoning prompts robust or fragile.
the dimensions that matter are structural, not logical
-
Do large language models reason symbolically or semantically?
Can LLMs follow explicit logical rules when those rules contradict their training knowledge? Testing whether reasoning operates independently of semantic associations reveals what computational mechanisms actually drive LLM multi-step inference.
same source batch: if reasoning is semantic not symbolic, logical validity of exemplars is irrelevant
-
Do reasoning traces need to be semantically correct?
Can models learn to solve problems from deliberately corrupted or irrelevant reasoning traces? This challenges assumptions about what makes intermediate tokens useful for learning.
convergent finding from training rather than prompting: invalid exemplars (this note) and corrupted training traces (that note) both preserve performance, confirming that logical content is dispensable and structure/scaffolding is the active ingredient
-
What do models actually learn from chain-of-thought training?
When models train on reasoning demonstrations, do they memorize content details or absorb reasoning structure? Testing with corrupted data reveals which aspects of CoT samples actually drive learning.
the structural explanation for why invalid logic still works: CoT gains come from structural coherence (step decomposition, scaffolding) not content correctness, so logically invalid exemplars provide the same structural benefits
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Invalid Logic, Equivalent Gains: The Bizarreness of Reasoning in Language Model Prompting
- Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
- When More is Less: Understanding Chain-of-Thought Length in LLMs
- Think Deep, Not Just Long: Measuring LLM Reasoning Effort via Deep-Thinking Tokens
- Mind Your Step (by Step): Chain-of-Thought can Reduce Performance on Tasks where Thinking Makes Humans Worse
- Deciphering the Factors Influencing the Efficacy of Chain-of-Thought: Probability, Memorization, and Noisy Reasoning
- Measuring Faithfulness in Chain-of-Thought Reasoning
- CoT is Not True Reasoning, It Is Just a Tight Constraint to Imitate: A Theory Perspective
Original note title
logically invalid cot prompts perform nearly as well as valid ones — valid reasoning is not the chief driver of cot gains