SYNTHESIS NOTE
Topics›Philosophy Subjectivity›this note

Can we predict where language models will fail?

Does characterizing the abstract computational problem an LLM solves—as a probability machine over sequences—let us predict which tasks it will struggle with systematically, before running experiments?

Synthesis note · 2026-05-18 · sourced from Philosophy Subjectivity

The "Levels of Analysis for LLMs" argument carries a specific empirical payoff: characterizing the abstract computational problem an LLM solves predicts where it will fail. The "embers of autoregression" line of work (McCoy et al.) is the worked example. By framing LLMs at Marr's computational level — as systems that have learned an autoregressive distribution over text — the researchers could derive in advance that tasks whose target response has low probability under the pretraining distribution would be systematically harder, even when the task itself is logically trivial.

The prediction is non-obvious. From a behavioral standpoint, you might expect difficulty to track task complexity. From the computational-level standpoint, you expect difficulty to track target probability, because the system is fundamentally a probability machine over sequences. Tasks like "write the alphabet backwards" or "count uppercase letters" can be logically simple but require generating sequences the pretraining distribution rarely supports. The framework predicted these would be hard before the experiments were run, and they were.

This is a working example of why a level-of-analysis approach is useful. Without it, the failure modes look like random capability gaps that need to be patched one by one. With it, the gaps look like predictable consequences of a particular kind of system, and they can be enumerated systematically by examining the computational characterization. The researcher who knows what problem the system is actually solving knows where to look for failure.

For interpretability research broadly, this is a template. Find the right computational-level characterization, derive its predictions about where the system should be brittle, and the brittle spots become a research program rather than an exception list.

Inquiring lines that read this note 126

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What compositional reasoning failures limit large language models despite scale? Why is hallucination an inevitable limitation of current language models? How do surface patterns enable correct outputs but reduce robustness? Is language model reasoning authentic and what causes models to reason? How do capability benchmark scores systematically misrepresent true model abilities? How do neural networks achieve compositional generalization at scale? Why don't LLMs reliably translate capability into accurate outputs? Do reasoning benchmarks predict model performance in long-horizon workflows? Do language models learn genuine understanding or just surface patterns? What enables genuine semantic understanding in language models? Why do token-level mechanisms matter for learning to reason? How much do training data properties shape model reasoning? How do multi-agent LLM systems fail distinctly compared to single agents? What articulatory and acoustic information does speech preserve that transcription destroys? Why do stronger reasoning capabilities create tradeoffs with instruction following? What role does sparsity play in model behavior and scaling decisions? Does encoded knowledge in language models actually influence their outputs? How does improved reasoning affect models' ability to acknowledge uncertainty? How do prompting refinements mask underlying biases and model frequency patterns? Can prompt-based context override biases that were embedded during pretraining? How effectively can language models perform reasoning, especially combined with symbolic methods? Do language models reason through causal mechanisms or semantic associations? How do prompt design choices influence model reasoning and performance? Do language models develop actual world models or merely task heuristics? Can compression size predict model complexity better than parameter count alone? What capability trade-offs arise from domain specialization through fine-tuning? Why do embedding systems fail to capture task-relevant relationships? What types of diversity prevent reasoning systems from collapsing? What training dynamics and scale trigger emergence of reasoning capabilities? Can mechanistic interpretability reliably guide practical model design choices? How do LLM judges' systematic biases affect alignment and evaluation outcomes? Can diffusion models match autoregressive performance on language generation tasks? Do language models reason like humans or mimic surface patterns? What makes imperfect LLM judges safe for optimization? Can harness architecture and protocols provide agent reliability without model scaling? Can reasoning scale in latent space without tokens?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 152 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

the computational level predicts where LLMs fail — embers of autoregression anticipated low-probability target failures