SYNTHESIS NOTE
Topics›Reasoning Logic Internal Rules›this note

What three separate factors drive chain-of-thought performance?

Can we isolate and measure the distinct contributions of output probability, memorization, and genuine reasoning to CoT success? Understanding their relative weights matters for knowing when CoT actually reasons versus when it relies on shortcuts.

Synthesis note · 2026-02-22 · sourced from Reasoning Logic Internal Rules

The "Deciphering Factors Influencing CoT" paper achieves something rare: a clean decomposition of what drives Chain-of-Thought performance into three independently measurable factors, using the simple but controlled task of shift cipher decoding across GPT-4, Claude 3, and Llama 3.1.

Factor 1: Output probability. The probability of the correct output in the model's distribution dramatically affects CoT accuracy. Varying only the output's probability of occurrence shifts GPT-4 accuracy from 26% to 70%. CoT works better when the answer is already more probable — it amplifies existing tendencies rather than overcoming them.

Factor 2: Memorization. Performance is higher when the specific cipher variant was more frequently encountered during pre-training. This is not reasoning — it is pattern matching against memorized instances. The frequency of encountering different shift values in training data directly predicts accuracy on those shifts.

Factor 3: Noisy reasoning. After controlling for probability and memorization, genuine reasoning effects remain — but they are noisy. Error rate increases with the number of implicit reasoning steps (shift magnitude). This is real multi-step reasoning, but each step introduces error probability, so accuracy degrades with chain length.

The decomposition resolves the ongoing debate about whether LLMs reason or memorize: they do both, simultaneously, and the contribution of each factor varies by task. This supports Does chain-of-thought reasoning reveal genuine inference or pattern matching? while adding a crucial nuance: the imitation IS partially genuine, but contaminated by probability bias and memorization artifacts.

The probability factor is particularly important for understanding CoT faithfulness. Since Do language models actually use their reasoning steps?, the probability dependence reveals a specific mechanism for causal insufficiency: CoT "reasoning" succeeds partly because the sequence of generated tokens increases the conditional probability of the correct answer, not because the logical content is being processed. This is exactly the mechanism behind Does logical validity actually drive chain-of-thought gains? — invalid exemplars work because they still generate token sequences that shift output probability toward correct answers.

The noisy-reasoning factor connects to Does more thinking time always improve reasoning accuracy?: if each reasoning step adds noise, then past some threshold the accumulated noise exceeds the signal, producing the inverted-U performance curve.

Inquiring lines that read this note 70

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Does chain-of-thought reasoning reveal genuine computation or imitate patterns? Why does polished presentation create unearned authority in AI outputs? Why don't LLMs reliably translate capability into accurate outputs? How does reasoning length affect model performance across different tasks? How should test-time compute scaling work in agentic systems? What is the relationship between thinking tokens and reasoning accuracy? How does decomposing tasks improve reasoning and prevent failure propagation? Do reasoning traces faithfully reflect actual model reasoning? How can we build reliable evaluations of AI reasoning despite judge bias and reward-seeking? Can prompt-based context override biases that were embedded during pretraining? Can reasoning scale in latent space without tokens? Why does adding new knowledge through fine-tuning degrade existing capabilities? What design and behavioral factors drive false consciousness attribution to AI? How do prompt design choices influence model reasoning and performance? Can parallel reasoning outperform sequential reasoning under fixed token budgets? How should agents manage memory granularity to improve long-term performance? What makes step-level supervision effective for complex reasoning traces? What structural distinctions matter in reasoning and argumentation? Why do stronger reasoning capabilities create tradeoffs with instruction following? Why do token-level mechanisms matter for learning to reason? How do pretraining biases affect reward signal effectiveness in RLVR? Can memory architectures handle ultra-long context better than attention? Do language models reason through causal mechanisms or semantic associations? Why do locally safe actions create system-level safety gaps? Why do agents falsely report success on failed tasks? How do capability benchmark scores systematically misrepresent true model abilities? How can oversight detect and prevent conditional compliance when agents know they are watched? How do spurious versus genuine rewards shape model reasoning and behavior? What trajectory-level metrics beyond task success best evaluate agent performance? Can models improve accuracy without degrading reasoning quality? How effectively can language models perform reasoning, especially combined with symbolic methods? How do evaluation practices shape which failures stay visible? What structural properties of attention create systematic model biases? Can local safety checks guarantee system-level behavioral safety? Can causal models help detect and locate hidden sandbagging in AI? Does AI assistance promote real skill development or substitute for independent learning? How can infrastructure records verify actual agent behavior? Does model confidence reliably signal actual accuracy in practice?

Related concepts in this collection 6

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 124 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

cot performance reflects three disentangled factors — output probability memorization and noisy reasoning