SYNTHESIS NOTE
Topics›Novel Architectures›this note

Can algorithms control LLM reasoning better than LLMs alone?

Explores whether embedding LLMs within algorithmic control flow—where programs manage state and context filtering—enables complex task decomposition beyond what LLMs achieve through self-managed reasoning chains.

Synthesis note · 2026-02-23 · sourced from Novel Architectures

LLM Programs embed an LLM within an algorithm rather than asking the LLM to be the algorithm. The critical design choice: instead of the LLM maintaining the current state of the program (its context), the LLM is presented with only step-specific prompt and context for each step. A classic computer program (Python) handles control flow, parsing of outputs, and augmentation of prompts for succeeding steps.

This is distinct from both Chain-of-Thought (where the LLM manages state through its token stream) and agentic frameworks (where the LLM decides what to do next). In LLM Programs, the algorithm structure is external and explicit, not learned or generated:

The key benefit is information hiding. By concealing information irrelevant to the current step, each LLM call focuses on an isolated subproblem whose results feed future calls. This addresses two fundamental limitations:

  1. Capability limits: Complex tasks that are currently too difficult because they require coordinating multiple reasoning steps
  2. Architectural constraints: The finite context window restricts processing to what fits within it

The approach recognizes the LLM as a limited general agent and avoids further training. Instead, the expected behavior is recursively deconstructed into simpler steps the LLM can perform to a sufficient degree.

This connects to Can modular cognitive tools unlock reasoning without training? — both decompose reasoning into modular operations. But LLM Programs are more structured: the control flow is predetermined by the algorithm, whereas cognitive tools are flexibly invoked. It also extends Does separating planning from execution improve reasoning accuracy? — the program IS the decomposer, and each LLM call IS the solver, with clean separation enforced by architecture rather than training.

Decomposed Prompting as the software library formalization: Decomposed Prompting (Khot et al., 2022) makes the software library analogy explicit. The decomposer defines a top-level program using interfaces to simpler sub-task functions. Sub-task handlers serve as "modular, debuggable, and upgradable implementations" — if a particular handler underperforms, it can be debugged in isolation, replaced with an alternative prompt or even a symbolic system (e.g., Elasticsearch), and plugged back in. This is more general than least-to-most prompting: it supports recursive decomposition, non-linear structures, and mixed neural-symbolic pipelines. The key architectural insight is that sub-task handlers are shared across tasks, creating a reusable prompt library — the closest existing analog to how software engineers build with functions. Source: Prompts Prompting.

Inquiring lines that read this note 159

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Why do persona simulations fail to predict authentic user behavior? Do language models reason like humans or mimic surface patterns? Can intelligent routing over smaller models outperform scaling a single large model? Why don't LLMs reliably translate capability into accurate outputs? Why do stronger reasoning capabilities create tradeoffs with instruction following? How do standardized protocols improve multi-agent coordination and reliability? What makes step-level supervision effective for complex reasoning traces? Can self-generated feedback reliably guide model training without ground truth? How effectively can language models perform reasoning, especially combined with symbolic methods? How can humans maintain meaningful oversight as AI systems become increasingly autonomous and complex? What causes reasoning models to fail or wander off track? How should designers communicate what AI systems truly are and can do? What reasoning architectures enable models to solve complex problems efficiently? Do knowledge graphs offer advantages over embeddings for multi-hop retrieval? Is language model reasoning authentic and what causes models to reason? How do multi-agent LLM systems fail distinctly compared to single agents? When should work require human-AI partnership versus full automation? How does decomposing tasks improve reasoning and prevent failure propagation? What capability trade-offs arise from domain specialization through fine-tuning? Why can recurrent transformers achieve reasoning capabilities that standard transformers cannot? Can brute-force automated research substitute for iterative depth and human research intuition? How should test-time compute scaling work in agentic systems? How should agents manage memory granularity to improve long-term performance? Should agents decouple planning from perception grounding for better performance? Why do locally safe actions create system-level safety gaps? Do language models reason through causal mechanisms or semantic associations? Do language models develop actual world models or merely task heuristics? Can parallel reasoning outperform sequential reasoning under fixed token budgets? How do prompting refinements mask underlying biases and model frequency patterns? Does RL create genuinely new reasoning capabilities or refine existing ones? Do reasoning benchmarks predict model performance in long-horizon workflows? Can memory architectures handle ultra-long context better than attention? Can compression size predict model complexity better than parameter count alone? How do prompt design choices influence model reasoning and performance? How should systems decide whether to retrieve or reason alone? Does chain-of-thought reasoning reveal genuine computation or imitate patterns? Do reasoning traces faithfully reflect actual model reasoning? How does misalignment propagate through agent communication networks? Should GUI agents use structured representations over raw visual input? What compositional reasoning failures limit large language models despite scale? Why is hallucination an inevitable limitation of current language models? Can local safety checks guarantee system-level behavioral safety? What execution architectures enable agents to most effectively use tools? Why do token-level mechanisms matter for learning to reason? How should agent systems validate and persist generated code artifacts? How does harness optimization generalize across different model architectures and domains? Can validator consensus certify semantic correctness beyond agreement? Do multi-agent systems introduce security vulnerabilities that single-agent architectures avoid? How does reasoning length affect model performance across different tasks? What makes imperfect LLM judges safe for optimization? Can single-point security defenses protect multi-agent systems from multi-step attacks? How do coordinated agents balance protocol compliance with reward maximization? Can welfare maximization and minority veto protection coexist? Can harness architecture and protocols provide agent reliability without model scaling? How do surface patterns enable correct outputs but reduce robustness?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 133 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

LLM programs decompose complex tasks into step-specific prompts within algorithmic control flow — hiding irrelevant context per step