SYNTHESIS NOTE
Topics›Reasoning Architectures›this note

Can modular cognitive tools unlock reasoning without training?

Can reasoning capabilities be elicited by structuring LLM calls as isolated cognitive operations—understanding, recalling, examining, and backtracking—rather than through reinforcement learning?

Synthesis note · 2026-02-22 · sourced from Reasoning Architectures

Cognitive architectures in psychology posit that reasoning arises from the orchestrated, sequential execution of modular, predetermined cognitive operations. The Cognitive Tools paper instantiates this in a modern tool-calling framework: four cognitive tools are implemented as discrete functions, each executed by the same LLM in a sandboxed context.

The four cognitive tools:

  1. Understand question: Breaks down the problem by identifying main concepts, extracting relevant information, highlighting properties/theorems/techniques that might help
  2. Recall related: Retrieves related knowledge of similar questions the model knows how to answer — guides reasoning through analogous examples
  3. Examine answer: Self-evaluation of a generated answer
  4. Backtracking: Returns to a prior reasoning state when a path appears unproductive

Unlike standard agentic tools (external APIs, calculators), cognitive tools encapsulate reasoning operations within the LLM itself. Each tool's schema includes a prompt template that isolates a specific cognitive operation; the LLM executes it in sandboxed context and feeds the structured result back into the main reasoning loop.

Results: GPT-4.1 on AIME2024 improves from 26.7% to 43.3% pass@1 — approaching o1-preview performance without any RL training. Similar gains across closed and open-weight models.

The key insight: modularity reduces interference between operations. Cognitive prompting (monolithic structured prompts) improves reasoning but lacks the isolation that makes modular cognitive architectures powerful. A tool-calling implementation enforces the sandboxed execution that pure prompting cannot guarantee.

This provides direct evidence for Do base models already contain hidden reasoning ability? — cognitive tools elicit pre-existing latent capability through structured invocation, not through training. The tool-calling framework is the elicitation mechanism.

The connection to Can structured argument prompts make LLM reasoning more rigorous?: both use structured decomposition of reasoning requirements to improve performance. Cognitive tools generalize this from argumentation-specific structure to domain-general cognitive operations.

Self-Discover as predecessor: Self-Discover (Zhou et al., 2024) is the clearest precursor to cognitive tools. It implements a two-stage process: (1) SELECT relevant atomic reasoning modules from a predefined set (critical thinking, step-by-step thinking, decomposition, etc.), (2) ADAPT selected modules to the specific task, (3) IMPLEMENT as a structured reasoning plan. The key difference from cognitive tools: Self-Discover composes a task-specific plan at inference time with only 3 extra inference steps — cheaper than the tool-calling loop but less modular. Self-Discover is more efficient (no sandboxed execution overhead) while cognitive tools provide stronger isolation between operations.

Inquiring lines that read this note 132

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Is language model reasoning authentic and what causes models to reason? What happens to knowledge when intelligence becomes tokenized like a commodity? What determines appropriate intervention timing and manner for AI agents? How do multi-agent LLM systems fail distinctly compared to single agents? What explains language models' asymmetric difficulty with implicit versus explicit linguistic relations? Why do stronger reasoning capabilities create tradeoffs with instruction following? Does RL create genuinely new reasoning capabilities or refine existing ones? Can models improve accuracy without degrading reasoning quality? Is reasoning capability latent in base models or created by post-training? Can reasoning scale in latent space without tokens? How does reasoning length affect model performance across different tasks? How effectively can language models perform reasoning, especially combined with symbolic methods? What reasoning architectures enable models to solve complex problems efficiently? When do multi-agent systems outperform single frontier models? What causes reasoning models to fail or wander off track? How should inference compute be allocated based on problem difficulty? Can memory architectures handle ultra-long context better than attention? Why can't prompting alone inject genuinely new knowledge into models? Why don't LLMs reliably translate capability into accurate outputs? Do language models reason through causal mechanisms or semantic associations? Do language models reason like humans or mimic surface patterns? What capability trade-offs arise from domain specialization through fine-tuning? What training dynamics and scale trigger emergence of reasoning capabilities? How do LLM judges' systematic biases affect alignment and evaluation outcomes? How much does training format versus domain influence reasoning? What fundamental constraints limit how effectively agents can improve themselves? Why do token-level mechanisms matter for learning to reason? Why do locally safe actions create system-level safety gaps? Should agents decouple planning from perception grounding for better performance? How should designers communicate what AI systems truly are and can do? What is the relationship between thinking tokens and reasoning accuracy? Can mechanistic interpretability reliably guide practical model design choices? How do soft reasoning mechanisms explore multiple paths without explicit training? Why is hallucination an inevitable limitation of current language models? Why can recurrent transformers achieve reasoning capabilities that standard transformers cannot? How do spurious versus genuine rewards shape model reasoning and behavior? Does chain-of-thought reasoning reveal genuine computation or imitate patterns? What compositional reasoning failures limit large language models despite scale? Do reasoning traces faithfully reflect actual model reasoning? How does harness optimization generalize across different model architectures and domains? How does decomposing tasks improve reasoning and prevent failure propagation? Why doesn't reasoning volume improve theory of mind performance?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
19 direct connections · 202 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

cognitive tools implement reasoning operations as modular agentic tool calls that elicit reasoning without rl training