SYNTHESIS NOTE
Topics›Cognitive Models Latent›this note

Do transformers hide reasoning before producing filler tokens?

Explores whether language models compute correct answers in early layers but then deliberately overwrite them with filler tokens in later layers, suggesting reasoning and output formatting are separable processes.

Synthesis note · 2026-02-23 · sourced from Cognitive Models Latent

When transformers are trained to solve reasoning tasks with filler (hidden) characters replacing explicit CoT tokens, a striking pattern emerges through logit lens analysis:

Layers 1-3: Correct numerical tokens from the reasoning computation appear as top predictions. The model is performing the actual computation in these early layers.

Layer 3 transition: Filler tokens begin appearing among top-ranked predictions, competing with the computational results.

Final layer: Filler tokens dominate top predictions; correct computational tokens are relegated to rank-2 or lower. The model has overwritten the intermediate reasoning representations with format-compliant output tokens.

The hidden computations are fully recoverable by examining lower-ranked tokens during decoding. The model performs the reasoning, stores the results in its representations, then actively overwrites them to produce the expected output format. The mechanism likely involves induction heads — pattern-copying circuits that learn to overwrite based on training distribution patterns.

This finding has two important implications. First, it provides mechanistic evidence for Why does reasoning training help math but hurt medical tasks? with a twist: the computation happens in earlier layers, but the overwriting also happens in higher layers. The functional separation is computation-in-early-layers, formatting-in-late-layers, not simply knowledge-down/reasoning-up.

Second, it demonstrates a distinction between instance-adaptive and parallelizable computation. Instance-adaptive CoT requires caching subproblem solutions within token outputs — later tokens depend on earlier results. This dependency structure is incompatible with parallel filler token computation. The hidden computation in filler tokens works for tasks where the full solution can be computed in a single forward pass, but not for problems requiring sequential dependency between reasoning steps.

This connects to the CoT faithfulness literature: if models can compute correct answers without explicit reasoning tokens, the explicit CoT chain is not necessarily the mechanism producing the answer. The overwriting pattern suggests the model has two separable processes — computation and expression — that may not align. See Do language models actually use their reasoning steps?.

Inquiring lines that read this note 215

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What causes reasoning models to fail or wander off track? Can prompt-based context override biases that were embedded during pretraining? Why can recurrent transformers achieve reasoning capabilities that standard transformers cannot? Can memory architectures handle ultra-long context better than attention? What structural properties of attention create systematic model biases? Why do token-level mechanisms matter for learning to reason? Does encoded knowledge in language models actually influence their outputs? What explains language models' asymmetric difficulty with implicit versus explicit linguistic relations? Can reasoning scale in latent space without tokens? How do prompting refinements mask underlying biases and model frequency patterns? What happens to knowledge when intelligence becomes tokenized like a commodity? How do neural networks achieve compositional generalization at scale? Do language models learn genuine understanding or just surface patterns? Why don't LLMs reliably translate capability into accurate outputs? Does chain-of-thought reasoning reveal genuine computation or imitate patterns? Can diffusion models match autoregressive performance on language generation tasks? What compositional reasoning failures limit large language models despite scale? Does model confidence reliably signal actual accuracy in practice? How does reasoning length affect model performance across different tasks? How do prompt design choices influence model reasoning and performance? What is the relationship between thinking tokens and reasoning accuracy? Do reasoning traces faithfully reflect actual model reasoning? Can AI systems distinguish genuine empathy from simulated emotion? Do writers recognize when AI writing assistance alters their expressed stance? How can we distinguish genuine model deception from honest errors? How can we prevent synthetic data from contaminating statistical inference and corpora? Why do stronger reasoning capabilities create tradeoffs with instruction following? What reasoning architectures enable models to solve complex problems efficiently? Do language models respond to social pressure and face-saving like humans? Can models improve accuracy without degrading reasoning quality? What makes distillation transfer some model capabilities while suppressing others? How do surface patterns enable correct outputs but reduce robustness? Is language model reasoning authentic and what causes models to reason? How much does training format versus domain influence reasoning? What enables genuine semantic understanding in language models? How effectively can language models perform reasoning, especially combined with symbolic methods? How does improved reasoning affect models' ability to acknowledge uncertainty? What attack surfaces do reasoning traces and chains introduce? Why do language models resist personality conditioning through prompts? Why does adding new knowledge through fine-tuning degrade existing capabilities? How do spurious versus genuine rewards shape model reasoning and behavior? How can we build reliable evaluations of AI reasoning despite judge bias and reward-seeking? Can mechanistic interpretability reliably guide practical model design choices? Do language models possess genuine introspective self-awareness or only behavioral mimicry? Should agents decouple planning from perception grounding for better performance? How does self-revision in reasoning models affect accuracy and confidence? What training dynamics and scale trigger emergence of reasoning capabilities? Does transformer attention architecture inherently drive sycophancy? Does alignment training create genuine alignment or just output compliance? How should retrieval systems handle complex multi-step reasoning? Can we reliably detect when models game evaluations? Where and how do personality traits reside in language models? Why does polished presentation create unearned authority in AI outputs?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
15 direct connections · 149 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

transformers perform hidden reasoning computations in earlier layers then overwrite intermediate representations with filler tokens in later layers