SYNTHESIS NOTE
Topics›Discourses›this note

Why do language models ignore information in their context?

Explores why language models sometimes override contextual information with prior training associations, and whether providing more context can solve this problem.

Synthesis note · 2026-02-21 · sourced from Discourses

The REMEDI paper names a specific failure mode: "failure of context integration." The example: an LM is prompted with a context establishing that Anita works in a law office, but when generating a continuation, the LM describes Anita as a nurse — overriding the contextual information with a prior association (names like Anita may statistically co-occur with certain occupations in training data).

This is a named, empirically documented failure mode, not a hypothetical. The failure occurs because the LM's parametric knowledge (compressed into weights from training) and its in-context information (the prompt) are not cleanly integrated. When they conflict, the parametric association can win.

The implication is important for how we think about context windows and RAG-style augmentation. Just providing information in context does not guarantee that a model will use it. If the information conflicts with strong prior associations, the prior may dominate — not because the model misread the context, but because context integration is not a lossless operation. The provided information gets processed through the same mechanisms that already have strong priors.

Fixing this requires causal intervention, not just better prompting: you need to modify the representations that carry the prior association, not just add more context on top of them. This is what REMEDI demonstrates — that adding a learned vector directly to entity representations can override the prior in a way that textual prompting cannot.

Inquiring lines that read this note 348

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What articulatory and acoustic information does speech preserve that transcription destroys? What compositional reasoning failures limit large language models despite scale? What determines appropriate intervention timing and manner for AI agents? How should conversational recommenders balance preference elicitation with direct recommendation? Do language models reason like humans or mimic surface patterns? Can prompt-based context override biases that were embedded during pretraining? What reasoning architectures enable models to solve complex problems efficiently? Does encoded knowledge in language models actually influence their outputs? What mechanisms preserve shared understanding in evolving conversations? How do surface patterns enable correct outputs but reduce robustness? What enables genuine semantic understanding in language models? What explains language models' asymmetric difficulty with implicit versus explicit linguistic relations? Do language models learn genuine understanding or just surface patterns? How do prompting refinements mask underlying biases and model frequency patterns? Why do LLM recommenders underperform collaborative filtering despite their capabilities? What training dynamics and scale trigger emergence of reasoning capabilities? Do structural constraints outperform deep architectures in recommendation systems? How do false presuppositions and sycophancy drive persistent false beliefs in models? How do training data properties determine the emergence of internal misalignment? What prevents conversational agents from taking initiative in dialogue? Can reasoning scale in latent space without tokens? Can compression size predict model complexity better than parameter count alone? Why do token-level mechanisms matter for learning to reason? Why can't prompting alone inject genuinely new knowledge into models? How do spurious versus genuine rewards shape model reasoning and behavior? How do LLM judges' systematic biases affect alignment and evaluation outcomes? Why do embedding systems fail to capture task-relevant relationships? Can diffusion models match autoregressive performance on language generation tasks? Can language models build genuine grounding through interaction? Why does adding new knowledge through fine-tuning degrade existing capabilities? How should systems decide whether to retrieve or reason alone? How should retrieval systems handle complex multi-step reasoning? Is language model reasoning authentic and what causes models to reason? Why do some clarifying approaches produce understanding while others just satisfy? What structural properties of attention create systematic model biases? How much do training data properties shape model reasoning? Can AI systems distinguish genuine empathy from simulated emotion? What capability trade-offs arise from domain specialization through fine-tuning? How well do AI systems understand human social norms? Does alignment training create genuine alignment or just output compliance? Can models improve accuracy without degrading reasoning quality? How does persona conditioning amplify demographic stereotyping and bias in models? How do prompt design choices influence model reasoning and performance? How should designers communicate what AI systems truly are and can do? Why do stronger reasoning capabilities create tradeoffs with instruction following? What causes retrieval-augmented generation systems to fail despite access to external knowledge? Can inoculation prompting prevent emergent misalignment after reward hacking? Can harness architecture and protocols provide agent reliability without model scaling? Why do language models resist personality conditioning through prompts? Does preference optimization systematically degrade conversational grounding in language models? Do language models respond to social pressure and face-saving like humans? Does abstract user knowledge outperform concrete interaction history in personalization? Should AI communication design follow human conversation norms or develop distinct machine-specific principles? Can local safety checks guarantee system-level behavioral safety? How do neural networks achieve compositional generalization at scale? Why is hallucination an inevitable limitation of current language models? Do knowledge graphs offer advantages over embeddings for multi-hop retrieval? How does improved reasoning affect models' ability to acknowledge uncertainty? Does model confidence reliably signal actual accuracy in practice? How can conversational agents maintain consistent personas across multi-turn dialogue? What attack surfaces do reasoning traces and chains introduce? How do neighboring agents influence whether others cooperate or collude? What causes reasoning models to fail or wander off track? Why don't LLMs reliably translate capability into accurate outputs? What factors drive AI persuasiveness and how can it be mitigated? Can memory architectures handle ultra-long context better than attention? Why can recurrent transformers achieve reasoning capabilities that standard transformers cannot? Do reasoning benchmarks predict model performance in long-horizon workflows? How does self-revision in reasoning models affect accuracy and confidence? Can mechanistic interpretability reliably guide practical model design choices? Can self-generated feedback reliably guide model training without ground truth? Do language models reason through causal mechanisms or semantic associations? How can we prevent synthetic data from contaminating statistical inference and corpora? What role does sparsity play in model behavior and scaling decisions? What training data selection strategies maximize generalization across difficulty levels? Should agents decouple planning from perception grounding for better performance? How should inference compute be allocated based on problem difficulty? Can causal models help detect and locate hidden sandbagging in AI?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
20 direct connections · 245 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

llm context integration fails when prior training associations override current context information