SYNTHESIS NOTE
Topics›Reasoning Methods CoT ToT›this note

Can minimal reasoning chains match full explanations?

Does removing all explanatory text from chain-of-thought reasoning preserve accuracy? This tests whether verbose intermediate steps are necessary for solving problems or just artifacts of how language models are trained.

Synthesis note · 2026-02-22 · sourced from Reasoning Methods CoT ToT

Chain of Draft (CoD) is a prompting strategy with a simple constraint: each intermediate reasoning step must be minimal — only the essential mathematical operation or logical transformation, with no explanation of what was done or why. The contrast with standard CoT is stark. Where CoT might produce six sentences to solve "20 - 12 = ?", CoD produces "20 - x = 12; x = 8."

The result: CoD matches or surpasses CoT accuracy across arithmetic reasoning, symbolic tasks, and commonsense tasks while using 7.6% of CoT's token count. The verbosity that CoT was assumed to require turns out to be unnecessary for the reasoning itself.

This challenges the implicit model underlying much test-time scaling work: that more tokens spent on reasoning generally produces better reasoning. The CoD finding suggests verbosity in CoT is a training artifact — LLMs are trained on human-written explanatory text, and CoT prompting induces that explanatory style even when the reasoning task only requires the critical operations. When you explicitly instruct minimal drafts, accuracy is preserved because the essential computation was never in the verbal explanation.

The mechanistic alignment with human note-taking behavior is telling: when humans do mental math, they jot down intermediate equations, not narrations of their own reasoning process. Standard CoT is asking LLMs to narrate their scratch work rather than write it.

This interacts with the Do reasoning traces actually cause correct answers? finding: if accuracy is preserved with 7.6% of the tokens, the other 92.4% was serving functions other than reasoning — explanatory style, human-readable documentation, or training-induced verbosity. The critical computation is localized in the minimal draft.

The practical implication for inference system design: token budget optimization should target verbose intermediate steps, not just final answer length. For tasks where CoD applies, you can run 13x more parallel chains under the same budget — combining the CoD efficiency advantage with Why does parallel reasoning outperform single chain thinking?.

Activation steering provides a mechanistic explanation for why CoD works. Can we steer reasoning toward brevity without retraining? shows that verbose and concise reasoning modes are geometrically separated in the residual stream. ASC (Activation-Steered Compression) extracts a steering vector from 50 paired examples and achieves 67% length reduction without retraining. This means CoD's prompting instruction ("keep each draft minimal") is a noisy way of pushing the model into the same activation region that the steering vector targets directly. The two methods are orthogonal and potentially combinable: CoD selects the concise region approximately through prompting, while ASC navigates to it precisely through activation intervention.

Inquiring lines that read this note 125

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Does chain-of-thought reasoning reveal genuine computation or imitate patterns? Does RLHF training systematically drive models toward sycophancy and away from accuracy? What causes reasoning models to fail or wander off track? How do prompting refinements mask underlying biases and model frequency patterns? What explains language models' asymmetric difficulty with implicit versus explicit linguistic relations? Does encoded knowledge in language models actually influence their outputs? Do reasoning traces faithfully reflect actual model reasoning? How do prompt design choices influence model reasoning and performance? How does reasoning length affect model performance across different tasks? How effectively can language models perform reasoning, especially combined with symbolic methods? Why do some clarifying approaches produce understanding while others just satisfy? Is reasoning capability latent in base models or created by post-training? What is the relationship between thinking tokens and reasoning accuracy? How should designers communicate what AI systems truly are and can do? Can reasoning scale in latent space without tokens? Can parallel reasoning outperform sequential reasoning under fixed token budgets? Is language model reasoning authentic and what causes models to reason? What makes distillation transfer some model capabilities while suppressing others? What reasoning architectures enable models to solve complex problems efficiently? Do language models reason through causal mechanisms or semantic associations? Can models improve accuracy without degrading reasoning quality? How should systems decide whether to retrieve or reason alone? Why do stronger reasoning capabilities create tradeoffs with instruction following? How do capability benchmark scores systematically misrepresent true model abilities? What safeguards enable trustworthy AI-assisted scientific peer review at scale? What enables genuine semantic understanding in language models?

Related concepts in this collection 9

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
23 direct connections · 225 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

concise intermediate reasoning chains match verbose cot accuracy with 7.6 percent of the tokens