Does reasoning ability actually degrade with longer inputs?
Explores whether modern language models can maintain reasoning performance when processing long contexts, and whether technical capacity translates to practical reasoning capability over extended text.
The FLenQA benchmark exposes a critical gap between technical context window capacity and actual reasoning capacity over long inputs. By embedding simple reasoning tasks (True/False questions requiring integration of two information pieces) within irrelevant padding text of varying lengths, the paper shows that reasoning accuracy drops from 0.92 to 0.68 at just 3000 tokens — far below any modern model's context window.
Three findings make this particularly concerning:
1. The degradation is task-agnostic. Regardless of whether padding text is similar or dissimilar to the reasoning content, and regardless of where the information pieces are embedded within the context, similar degradation trends appear. The failure is not about content interference but about attention dilution over length.
2. Next-word prediction performance is uncorrelated with reasoning performance. Models that maintain strong perplexity on long inputs still fail at reasoning over those inputs. This means language modeling benchmarks on long contexts are misleading indicators of actual long-context utility — a model can "understand" the text (predict tokens well) while failing to reason over it.
3. CoT does not mitigate proportionally. Chain-of-thought prompting increases accuracy roughly uniformly across context lengths but does not close the length-induced gap. The degradation persists under CoT because the bottleneck is in information retrieval from context, not in reasoning over retrieved information.
This is a complementary mechanism to Why do language models fail at temporal reasoning in complex tasks?. That failure is about task complexity; this is about input noise. Together they define a two-dimensional reliability surface: reasoning degrades with both task complexity AND input length, and the two dimensions are independent.
The implication for RAG systems is direct: retrieved documents add to input length, and if that length includes irrelevant passages (as it typically does), reasoning over the retrieved content degrades even when the relevant information is present. Since Why does vanilla RAG produce shallow and redundant results?, the length degradation explains part of why static retrieval fails — more retrieved documents means more padding means worse reasoning.
A complementary training-time finding complicates this picture. "Longer Context, Deeper Thinking" (2025) shows that models with stronger long-context capacity (128k vs 32k) consistently achieve higher accuracy on mathematical reasoning benchmarks (MATH500 and AIME) — even when test-time inputs are short. Long-context training benefits reasoning as a foundation, not just for processing long inputs. The implication: the inference-time degradation documented in this note coexists with a training-time benefit. Models trained on longer contexts develop better reasoning foundations, but at inference time, longer inputs still degrade performance. The two findings are compatible: long-context training may improve the base reasoning capability, while inference-time input length introduces the noise and distraction effects that degrade it. Source: Arxiv/Evaluations.
Inquiring lines that read this note 164
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Does AI assistance promote real skill development or substitute for independent learning? What enables genuine semantic understanding in language models?- How do readers selectively hold frame-related words in mind?
- What cognitive abilities distinguish metalinguistic analysis from language use?
- Can input augmentation and rephrasing compensate for smaller model limitations?
- Does filtering passages before generation improve large model answer quality?
- How do behavioral differentiation and paraphrase stability trade against accuracy?
- How do language models treat injected evidence as shared background knowledge?
- How does the knowing-doing gap widen as tasks become more complex?
- Do reasoning models trade instruction following for deliberative capability?
- Does reasoning structure match explicit versus implicit task demands?
- How does scaling reasoning capability actually reduce instruction-following ability?
- What changes when reasoning models adopt trajectory-response output formats?
- Do shorter reasoning chains maintain instruction adherence better than longer ones?
- Does sentence-level granularity capture enough structure for complex reasoning tasks?
- What makes a background condition relevant to a specific reasoning task?
- Where do humans and language models actually diverge in reasoning ability?
- Can marginal hints integrate better into reasoning than comprehensive explanations?
- Why do models automatically adjust reasoning length to problem difficulty?
- Why does explicit reasoning degrade passage reranking performance?
- What causes snowball errors to accumulate across reasoning steps in language models?
- Why does extended reasoning fail for search and knowledge retrieval tasks?
- Can long-context models handle compositional reasoning requiring structured logic?
- Why do reasoning models fail when input length increases even below context limits?
- Can simple structure perturbations reliably expose memorization in reasoning models?
- Can cognitive scaffolding replace tool-based reasoning augmentation in language models?
- What causes reasoning quality to degrade during long research tasks?
- Can bounded workspaces prevent overthinking better than summarization alone?
- Why might rationales that predict common text patterns fail on hard novel reasoning?
- When is numeric computation the real bottleneck versus reasoning depth?
- Is reasoning failure caused by task complexity or training distribution gaps?
- How does recombining partial trajectories maintain coherence in natural language reasoning?
- Can weak models reason better when freed from cognitive load by structure?
- Does irrelevant content degrade reasoning even when it fits the context window?
- Can manipulative prompts reduce reasoning model accuracy without fine-tuning?
- Does irrelevant context degrade reasoning even within model context limits?
- Can structured prompts reduce reasoning steps while improving financial accuracy?
- Can operationalizing theory into prompt structure improve reasoning more than theory itself?
- How does SONAR embedding quality affect downstream reasoning accuracy?
- What distinguishes genuine reasoning activation from memorization-assisted answer recall?
- How much does pre-training frequency predict reasoning task performance?
- Does model scaling improve knowledge storage faster than reasoning ability?
- How can entailment benchmarks separate genuine reasoning from memorization effects?
- Can dataset design systematically expand reasoning graph diameter?
- Why do reasoning tasks improve more than retrieval from lookup memory?
- What kinds of reasoning tasks reveal the ceiling of text-only training?
- How does evaluation setting affect measured reasoning capabilities in language models?
- Can latent reasoning architectures work as retrofits to existing models?
- Why might latent reasoning capture types of thinking that verbalized CoT cannot?
- Can latent reasoning mechanisms and recursive tracking mechanisms be combined effectively?
- Can context compression preserve what matters without introducing bias?
- How does reducing activation precision further extend context length?
- What happens to anaphoric reference when context exceeds the window?
- Why do longer context windows alone fail to capture temporal dynamics in dialogue?
- Why does long-form generation need different retrieval than factoid questions?
- Why do longer queries benefit less from clarification questions?
- When does long-context LLM reasoning fail where structured retrieval succeeds?
- Can long-context readers handle compositional tasks or just semantic search?
- Why do fixed-size document chunks break complex procedural question answering?
- What makes active reasoning through dialogue harder than passive reasoning?
- How does era sensitivity in legal cases compound with context length failures?
- When should an LLM engage extended reasoning versus responding directly?
- Could real-time search systems avoid era sensitivity in legal reasoning?
- What prompting strategies most effectively boost long-context LLM performance on retrieval?
- Can derivational traces be distinguished from stylistic mimicry of reasoning?
- Why do correct reasoning traces in language models tend to be shorter?
- Does explicit reasoning help or hurt tasks requiring continuous nuanced judgment?
- Can extended reasoning training capture individual strategic thinking styles?
- Does more thinking always help large language models or sometimes hurt?
- How does random walk length control reasoning complexity in question generation?
- Why do longer reasoning chains signal hesitation rather than depth?
- What structural properties define effective long chain-of-thought reasoning?
- How do smaller models respond to longer reflection prompts?
- Why does chain-of-thought prompting fail to fix length-induced reasoning degradation?
- Can models trained on longer contexts develop better fundamental reasoning abilities?
- How do longer reasoning chains create vulnerability to attacks?
- Why do longer reasoning chains correlate with lower accuracy in o1-like models?
- How much reasoning depth do we actually need for most real-world tasks?
- Can minimal reasoning steps match verbose reasoning accuracy?
- Does more thinking always improve language model accuracy?
- Does chain-of-thought accuracy degrade with longer reasoning traces?
- Why do thinking models execute longer tasks than standard language models?
- What happens to long-tail reasoning when AI assists public deliberation?
- Why does reasoning performance degrade as input length increases?
- Do earlier errors in long tasks increase the likelihood of future mistakes?
- Does longer reasoning always improve model accuracy on complex tasks?
- Do tokens beyond a critical threshold actually improve reasoning quality?
- Does thinking-token overuse actually degrade reasoning accuracy in practice?
- How does chain-of-thought length affect attention to constraint tokens?
- Why does representation recycling of MI-peak tokens improve reasoning accuracy?
- How much does schema bloat actually degrade reasoning in large language models?
- Why do concise reasoning chains match verbose chain-of-thought token efficiency?
- Can concise reasoning traces match verbose explanation accuracy?
- Why do temporal reasoning patterns matter more than final answers?
- Why does concise reasoning maintain accuracy with far fewer tokens?
- What quality filters distinguish useful reasoning enrichment from shallow repetition?
- Why does training data format shape reasoning strategy more than domain content?
- Does training data format shape reasoning strategy more than domain content?
- How much does training data format influence reasoning strategy versus domain content?
- Why do language models fail at pronouns across distant segments?
- Why do language models fail at coreference across long contexts?
- Why do large language models fail at temporal reasoning in complex legal cases?
- What neuroscience evidence suggests language networks are not optimized for reasoning?
- How does the distance between natural language and formal notation affect translation accuracy?
- Do pretrained language models carry reusable computational scaffolding for length handling?
- Why do long-context language models struggle with compositional reasoning tasks?
- Can autoformalisation from natural language preserve semantic accuracy?
- Are newer larger language models actually worse at faithful summarization?
- Can episodic and semantic memory improve long-horizon task reasoning?
- How does separating local and global context dependencies affect long-context performance?
- Does recurrent memory or gist compression work better for ultra-long context?
- Can recurrent state mechanisms process longer sequences than attention-based working memory approaches?
- How do adaptive memory modules compare to feedback-based working memory for long context?
- How does externalized state affect the long-context bottleneck in language models?
- How do recurrent memory systems handle ultra-long context differently than attention?
- Does more inference compute help reasoning models match specialized domain performance?
- How should iterative research tasks limit context per reasoning turn?
- What is the optimal balance between search rounds and reasoning depth per round?
- How should reasoning prompts adapt based on question complexity and type?
- How do input length and context size separately affect reasoning quality?
- How do neural memory modules extend context length beyond attention limits?
- Why does attention quality degrade as context length increases?
- Why does attention concentrate on the first 25% of long input sequences?
- How does context complexity affect LLM performance on temporal reasoning tasks?
- What makes deductive reasoning so brittle in language models overall?
- How does structural complexity in sentences degrade LLM reasoning systematically?
- Why do format and structure matter more than actual content in reasoning?
- Why does premise ordering shift syllogistic reasoning performance by over 30 percent?
- Can language models perform genuine symbolic reasoning without semantic grounding?
- Can benchmark performance distinguish surface from structural linguistic knowledge?
- Can language models reason without relying on surface level pattern matching?
- Can adding more words to a passage actually interfere with meaning?
- Does high knowledge density in text reduce user motivation to read more?
- What makes specific-facet questions outperform generic need-rephrasing requests?
- Why does document perplexity stay low while question-answering accuracy drops?
- How do retrieval heads interact with layer-level separation of knowledge and reasoning?
- How does context length affect retrieval quality in modernized BERT architectures?
- How does evidence retrieval affect compositional reasoning in language models?
- Do distributed relational tasks consistently underperform local classification across NLP domains?
- Should long-context evaluation measure the coupled system?
- Do base models truly possess latent reasoning capability?
- Can auxiliary modules preserve reasoning without catastrophic forgetting?
- Does sequence length affect sparsity tolerance the same way across task types?
- What makes sparse attention more reliable for long-context retrieval?
Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Why do language models fail at temporal reasoning in complex tasks?
Language models correctly answer simple temporal questions but produce logically impossible timelines in complex legal documents. This explores what task features trigger reasoning failures and whether the competence is genuinely lost or masked by surface-level patterns.
complementary failure axis: task complexity vs input length
-
Why does vanilla RAG produce shallow and redundant results?
Standard RAG systems get stuck in a single semantic neighborhood because their initial query determines what documents are discoverable. The question asks whether fixed retrieval strategies fundamentally limit knowledge depth compared to iterative exploration.
RAG retrieval adds length; length degrades reasoning
-
Does more thinking time actually improve LLM reasoning?
The intuition that extended thinking helps LLMs reason better seems obvious, but what does the empirical data actually show when we test it directly?
another dimension where "more" (tokens) ≠ "better" (reasoning)
-
Can long-context models resolve retriever-reader imbalance?
Traditional RAG systems force retrievers to find precise passages because readers had small context windows. Do modern long-context LLMs change what architecture makes sense?
challenges the long-context solution: reader burden increases with length but reasoning degrades
-
Do vector embeddings actually measure task relevance?
Vector embeddings rank semantic similarity, but RAG systems need topical relevance. When these diverge—as with king/queen versus king/ruler—does similarity-based retrieval fail in production?
compounds the length problem: semantic retrieval returns associated-but-irrelevant documents, creating exactly the irrelevant padding that FLenQA shows degrades reasoning; imprecise retrieval directly produces the input-length degradation
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models
- Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in Reasoning
- Think Deep, Not Just Long: Measuring LLM Reasoning Effort via Deep-Thinking Tokens
- Self-Guided Test-Time Training for Long-Context LLMs
- When More is Less: Understanding Chain-of-Thought Length in LLMs
- On the Reasoning Capacity of AI Models and How to Quantify It
- Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs
- Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
Original note title
reasoning performance degrades with input length even far below context window limits