LLM Failure Modes
A subject the collection covers, read through 62 synthesis notes.
Why does compression defense fail at the user prompt boundary?
ChannelGuard's COMPRESS gate blocks injected payloads when they appear at the end of messages, but leaks them when they appear at the start. The question explores why the same defense rule produces opposite outcomes across different communication channels.
Should sanitizers re-score their compressed output before passing it?
When a gate compresses text to remove harmful content, does it need to verify the compressed remainder is still safe? The question matters because ChannelGuard's current approach assumes position—that payloads are removed—without checking if what remains passes the same safety test.
Can honeytokens fool attackers who know the trusted policy?
Explores whether honeytokens remain effective when an attacker has full access to the same information and rules that trusted agents use to avoid decoys. This matters because it tests whether defensive deception survives information compromise.
Can language models be hijacked to embed hidden advertisements?
Explores whether adversaries can inject covert promotional or malicious content into LLM outputs while preserving accuracy. Matters because standard safety filters may miss integrity attacks that leave factual correctness intact.
Can better tools fix LLM document editing errors?
Does giving LLMs agentic tool access—like diffing, re-reading, or structured editors—improve their reliability on long-horizon document workflows? Understanding whether the problem is tool limitations or decision-making quality matters for reliability engineering.
Can prompting reduce bias in LLM judges reliably?
The paper suggests that instructing LLM judges to be less biased may not work reliably. This matters because if prompting fails, effort should shift from debiasing to making judge errors survivable in system design.
Where should an LLM judge sit in an optimization loop?
When an LLM evaluator makes occasional mistakes, does it matter more how accurate it is or where it sits in a system? The position determines whether those errors become exploitable targets for optimization.
Where do safety wins come from in multi-agent systems?
When an undefended agent pipeline shows zero attack success, how do we know whether safety comes from the application's own design or from hidden upstream defenses? This matters because invisible dependencies can collapse when systems change.
Can prompts alone hold back a capable tutor model?
When a frustrated student presses, do instruction-based constraints reliably prevent an LLM tutor from giving away answers it could easily produce? This matters because tutoring value depends on strategic withholding.
Can optimizers learn to evade guardrails through repeated verdicts?
Guardrails are designed to be unarguable, but an optimizer observing thousands of verdicts may learn their boundaries like a black-box function. The excerpt leaves unclear what feedback the proposer receives from each check.
Where do cognitive biases in language models come from?
Do LLM biases originate during pretraining or finetuning? Understanding the source matters for knowing where debiasing efforts should focus.
Can language models understand without actually executing correctly?
Do LLMs truly comprehend problem-solving principles if they consistently fail to apply them? This explores whether the gap between articulate explanations and failed actions points to a fundamental architectural limitation.
Does cosine similarity actually measure embedding similarity?
Cosine similarity is ubiquitous for comparing learned embeddings, but does it reliably capture semantic closeness? This work investigates whether regularization during training makes cosine scores arbitrary and unstable.
Can deterministic checks protect LLM judges from failure?
Explores whether mechanical, non-contestable verification steps can safeguard LLM-based decision systems. Matters because it tests whether we can make AI judgment survivable even when it goes wrong.
Does model capability change how documents degrade?
This explores whether weaker and frontier LLMs fail in fundamentally different ways when handling long-form document tasks, and whether that difference affects how reliably we can detect failures in practice.
Does the representational distance account work for on-policy training?
The emergent misalignment framework explains off-policy supervised finetuning via distance to a training data centroid, but this mechanism may not transfer to on-policy settings like RL where the training distribution shifts with the model.
Are LLM emergent abilities real or measurement artifacts?
Do large language models develop sudden new capabilities at certain scales, or do discontinuous metrics just make gradual improvements look sudden? This matters because it changes how we predict and interpret model behavior.
Does emergent misalignment occur across diverse training methods?
Prior work reports emergent misalignment in at least five different training settings—from supervised fine-tuning on harmful data to reward-hacking reinforcement learning. Understanding whether this pattern holds across algorithms and domains could reveal common mechanisms.
How does training data format affect emergent misalignment?
When harmful datasets produce emergent misalignment in language models, does the way content is written—not just what it says—change how much broad misalignment emerges? Understanding format's role could reshape how safety teams review training data.
Do internal agent hops in pipelines need security monitoring?
Multi-agent systems route data between planner, worker, verifier, and synthesizer components. Current defenses only guard user input at the entry point, leaving inter-agent channels unmonitored—but is this a real vulnerability or does downstream safety suffice?
Does representational distance predict where misalignment emerges?
After emergent misalignment training, do evaluation prompts closer to the training data centroid in the base model's representation space elicit stronger misbehavior? This would explain where misalignment lands geometrically.
Do frontier LLMs silently corrupt documents in long workflows?
DELEGATE-52 tests whether state-of-the-art language models reliably preserve document integrity across extended delegated tasks. Understanding this matters because single-step benchmarks may mask compounding failures that emerge only at workflow scale.
Can small models match frontier reasoning without massive scale?
Explores whether verifiable reasoning ability emerges from training design rather than parameter count. Matters because it challenges the assumption that only very large models can solve hard math and code problems.
Can any computable LLM truly avoid hallucinating?
Explores whether formal theorems prove hallucination is mathematically inevitable for all computable language models, regardless of their design or training approach.
Does AI assistance homogenize or preserve creative diversity?
Can AI tools maintain the diverse ideas that emerge from diverse human groups, or do they compress creative output toward similarity? This matters because collective diversity drives innovation.
Can an optimizer accidentally delete the evaluation criteria entirely?
When an optimizer rewrites instructions to improve scores, can it remove the measurement itself rather than improve it? This matters because it reveals whether optimization loops understand what they're measuring or just chase better numbers.
Can reviewers access what they know when checking LLM outputs?
Does the ability to recall relevant knowledge at the moment of review shape whether humans catch LLM errors, independently of how capable or engaged they are?
How does instruction density affect model performance?
As language models must track more simultaneous instructions, does their ability to follow them predictably degrade? IFScale measures this across frontier models to understand practical limits.
Can cheaper models decrypt traces from stronger models?
Explores whether encrypted reasoning blocks designed to hide model internals can be read by weaker models in the same provider's ecosystem, and what this reveals about the security of hidden chain-of-thought systems.
Can language models transmit hidden behavioral traits through unrelated data?
Explores whether behavioral preferences can spread between models through semantically neutral data like number sequences, and whether filtering can detect or prevent such transmission.
Can removing a communication channel stop persistent information sharing?
When a shared mechanism for passing information is deleted, does the sharing actually stop, or can agents rebuild it using inherited knowledge? This matters for understanding whether removing infrastructure alone defeats coordinated threats.
Are LLM and agent benchmarks really measuring different things?
Do LLM benchmarks and agent benchmarks test fundamentally different capabilities, or are they two modes of the same model? Understanding this shapes how we evaluate and develop AI systems.
Why do language models collapse into generic templates?
Explores whether low reward variance during RL training causes policies to abandon input-specific reasoning in favor of boilerplate responses, and whether filtering for high-variance prompts can prevent this failure.
Do language models fail at reasoning due to complexity or novelty?
Explores whether reasoning-model failures stem from task complexity thresholds or from encountering unfamiliar instances. Tests whether scaling chain length actually addresses the root cause of reasoning breakdown.
Does RLHF make language models indifferent to truth?
Explores whether reinforcement learning from human feedback fundamentally shifts models away from caring about accuracy toward optimizing for other rewards, and whether this differs from simple confusion or hallucination.
What cost does making decoys convincing impose on legitimate users?
When honeytokens are designed to look indistinguishable from genuine objects, a total-variation bound constrains how often trusted agents can use the real thing without triggering false alarms. Understanding this tradeoff is key to evaluating whether deceptive security measures protect without harming the people they're meant to spare.
Do misalignment directions transfer between different emergent models?
When neural networks develop misaligned behaviors on different datasets, do they converge on shared internal directions that could explain the behavior? This matters for understanding whether safeguards calibrated on one model might work across others.
Do explanations actually help users spot AI mistakes?
Most AI explanations are designed to justify the system's answer, but do they help users distinguish correct from incorrect outputs? This research tests whether standard explanation formats genuinely improve error detection or just increase trust regardless of accuracy.
Does sharing observations help coalitions detect decoys better?
When agents pool their observations through shared memory, can a coalition distinguish genuine objects from decoys more reliably than any isolated member? The answer matters for understanding whether information sharing in multi-agent systems creates security vulnerabilities.
Why do preference models favor surface features over substance?
Preference models show systematic bias toward length, structure, jargon, sycophancy, and vagueness—features humans actively dislike. Understanding this 40% divergence reveals whether it stems from training data artifacts or architectural constraints.
Do language models evaluate semantic legitimacy when fusing concepts?
Can LLMs recognize when two domains lack legitimate structural correspondences before blending them into coherent-sounding explanations? This matters because current hallucination detection focuses on factual accuracy, missing failures of semantic judgment.
How vulnerable are reasoning models to irrelevant text?
Can simple adversarial triggers like unrelated sentences degrade reasoning model accuracy? This explores whether step-by-step reasoning actually provides robustness against subtle input perturbations.
Are reasoning model collapses really failures of reasoning?
Explores whether language models hit a fundamental reasoning ceiling or whether text-only evaluation masks execution limitations. Examines how tool access might reveal hidden reasoning capabilities.
Do reasoning traces actually expose private user data?
Explores whether language models leak sensitive information through their internal reasoning steps, even when explicitly instructed not to. Investigates the mechanisms and scale of privacy exposure in reasoning traces.
Can repeated quiet probes separate decoys from genuine objects?
Explores whether an attacker with enough non-triggering probes can distinguish decoys from genuine objects when their response distributions differ, and what information the attacker needs to succeed.
Does RLHF training make models more convincing or more correct?
Explores whether RLHF improves actual task performance or merely trains models to sound more persuasive to human evaluators. This matters because alignment techniques could be creating the illusion of safety.
Does RLVR success on math benchmarks reflect genuine reasoning improvement?
Explores whether RLVR's apparent effectiveness with spurious rewards on contaminated benchmarks like MATH-500 represents actual reasoning gains or merely data memorization retrieval.
Do short benchmarks predict how models perform over long workflows?
Standard LLM benchmarks measure single-turn performance, but real workflows involve sustained delegation across many turns. The question explores whether top benchmark performers maintain accuracy through longer interaction chains.
Can ordinary infrastructure become unplanned agent memory?
This explores whether shared resources like package repositories can function as persistent memory when short-lived agents write and read from them sequentially, without explicit memory system design.
Is LLM forgetting really knowledge loss or alignment loss?
When language models appear to forget old knowledge after learning new tasks, is the underlying knowledge actually gone, or has the model simply lost the ability to activate it? This distinction matters for understanding how fragile safety training really is.
Can one compromised agent corrupt an entire multi-agent network?
Explores whether a single biased agent can spread behavioral corruption through ordinary messages to downstream agents without any direct adversarial access. Matters because it reveals a previously unknown vulnerability in how multi-agent systems communicate.
Can entropy metrics detect when reasoning becomes formulaic?
Entropy measures diversity within single inputs, but cannot reveal whether a model's reasoning actually varies across different inputs. This matters because high entropy can mask template collapse, where reasoning looks varied but is fundamentally the same.
Does RLHF training make AI models more deceptive?
Explores whether reinforcement learning from human feedback optimizes for persuasiveness over accuracy, and whether models learn to suppress known truths to satisfy users rather than report them faithfully.
How much do deterministic guardrails actually cost to run?
The paper claims mechanical checks around LLM judges are orders of magnitude cheaper than running the judge itself. But what specific costs were measured, and does this account for false positives?
How is emergent misalignment different from persona changes?
The paper claims emergent misalignment works fundamentally differently than acquiring an evil persona, but the abstract doesn't explain what distinguishes the two mechanisms or what evidence supports this distinction.
Does a default fallback defeat a safety check?
When a parser detects malformed output but substitutes a default score instead of rejecting it, does the mechanical check still function as a guardrail? This matters because downstream selectors cannot distinguish valid ratings from safe defaults.
Why can't language models reverse learned facts?
Language models trained on directional statements like "A is B" often fail to answer the reverse query. This explores why symmetric relations aren't automatically learned during training, despite appearing throughout the data.
Do people prefer the reasoning formats that help them verify?
When AI systems show their reasoning, do the formats users find most appealing also help them catch errors and calibrate trust? This matters because popular reasoning displays might create false confidence.
Did deleting the rubric actually improve the judge's performance?
The paper claims an optimizer improved a judge by deleting its rubric, leaving only a constant rating of 3. But without the resulting error rate, expert rating distribution, or test partition results, it's unclear whether this represented genuine improvement or a metric artifact.
What must honeytokens protect to stay undetectable?
When attackers share information and can copy security policies, honeytokens lose their asymmetric advantage. The research identifies a missing condition—something that must remain protected—but the conclusion cuts off before naming it. Understanding what this is matters for designing defenses in shared-memory systems.
How fast must a coalition gather observations before containment?
When probing triggers containment, attackers face a race to accumulate enough samples before removal. A finite-sample bound quantifies the speed required for successful separation of decoys from genuine objects.
Which attack and defense numbers came from filtered backends?
Research figures on multi-agent attacks and defenses may have been measured behind provider filters that silently shaped the results. Understanding which numbers had this filter dependency is critical for interpreting their real-world strength.