Does model capability change how documents degrade?
This explores whether weaker and frontier LLMs fail in fundamentally different ways when handling long-form document tasks, and whether that difference affects how reliably we can detect failures in practice.
DELEGATE-52 surfaces an under-discussed asymmetry in how LLM document degradation looks at different capability tiers. Weaker models fail loudly: they delete content. The document gets visibly shorter, sections disappear, structure breaks. A reviewer notices.
Frontier models fail quietly. Their degradation comes from corruption of existing content — values flipped, references rewritten, edits applied in the wrong place — producing documents that look intact at a glance but contain accumulated drift. The corruption mode is more dangerous than the deletion mode precisely because it preserves the surface signal of competence. The thing that looks like a successful workflow output is the thing that has silently drifted.
This matters for adoption. The "frontier models are reliable" intuition is built from short-interaction benchmarks where the corruption mechanism barely activates. At workflow scale — the regime where delegation is actually useful — the failure changes character, and the qualitative shift toward harder-to-detect failures means that improvements in raw capability can degrade overall workflow reliability if review effort is held constant.
The implication for delegated-AI design is that capability improvements at the frontier need to be paired with detection mechanisms that target corruption-style errors, not just deletion-style errors. Diff review, document-state checksums, and constraint validators become more important as models get better, not less.
Inquiring lines that read this note 70
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What compositional reasoning failures limit large language models despite scale? How do evaluation practices shape which failures stay visible?- What makes diverse failure modes more informative than single failure examples?
- Why do frontier model failures in document editing go undetected by users?
- What baseline evidence distinguishes amplification from unchanged failure rates?
- Can reliable failure detection prevent optimization pressure against detectors?
- Why do quiet failures reach deployment scale more often than loud ones?
- How do workflows normalize and hide errors before they become visible hazards?
- How does laboratory generalization evidence connect to deployment failure modes?
- What makes intermediate primitives matter more than final code execution success?
- What does it mean for errors to remain visible, contestable, and recoverable?
- What makes a model's errors visible and contestable to users?
- What distinguishes domain-specific failure modes from general model limitations?
- What makes some model capabilities reliable while others remain brittle?
- Why do frontier models corrupt more documents than weaker models during workflows?
- How does workflow scale change the failure modes of frontier models?
- How does model tier affect whether errors delete or corrupt document content?
- Why do frontier models corrupt documents while weaker models delete them?
- Can semantic audit layers attribute failure mechanisms to infrastructure-level state changes?
- Which failure modes dominate when models handle underspecified requests?
- What causes silent document corruption in long LLM workflows?
- What barriers prevent experts from specifying concepts for LLM extraction?
- Can lightweight verification methods help experts trust LLM outputs?
- Does LLM miscalibration cause failures in clinical information extraction?
- How long does retrievability support error detection across repeated LLM use?
- What levels of understanding about LLM knowledge representation can automated systems reliably extract?
- Can review effort alone keep pace with frontier model degradation?
- How much of weak-to-strong performance gaps reflect presentation rather than capability?
- Did AIDE2's rewrites solve problems on a human checklist or search artifacts?
- Does held-out validation prevent skill document edits from drifting or accumulating harm?
- Why does embedding research tools in coding assistants improve reliability?
- How do artifact families differ in matching verification scope to repair capability?
- How can anchored records fail authenticity while passing integrity checks?
- Do infrastructure event records alone suffice to distinguish different failure mechanisms?
- Does semantic validity across a quorum require new property definitions?
- How can we detect when protocol-compliant validators certify semantically incorrect states?
- Which shared channels cause the strongest correlated validator failures?
- What counts as full capability recovery versus partial restoration?
- How much does workflow architecture matter versus raw model capability?
- What happens to a commitment when its bound content must be deleted?
- Does the paper treat storage traces as addressed messages or unmarked traces?
- Why do mid-tier models benefit most from memorized harness fixes?
- What should an external contract for model improvement actually contain?
- Which frontier LLM models generate more misaligned emails than others?
- Which frontier LLM models generated the most misaligned emails?
- Can workflow-level validation reconstruct the global risk context that no single step holds?
- How much safety burden shifts between provider filters and model alignment in rerouted requests?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Do frontier LLMs silently corrupt documents in long workflows?
DELEGATE-52 tests whether state-of-the-art language models reliably preserve document integrity across extended delegated tasks. Understanding this matters because single-step benchmarks may mask compounding failures that emerge only at workflow scale.
same paper, the parent claim
-
Can better tools fix LLM document editing errors?
Does giving LLMs agentic tool access—like diffing, re-reading, or structured editors—improve their reliability on long-horizon document workflows? Understanding whether the problem is tool limitations or decision-making quality matters for reliability engineering.
same paper, why naive tool-use does not fix this
-
How does AI-generated false experience differ linguistically from human deception?
When AI writes about experiences it never had, does it leave distinct linguistic traces that differ measurably from intentional human lies? Understanding these differences could reveal how AI falsity is fundamentally different in structure.
adjacent: another mode of unfalsified-looking falsity
-
What makes quietly failing systems more dangerous than obvious ones?
Systems with obvious failures get caught and dropped before scale. But what conditions make a subtly failing system persist and spread? Why might that be worse?
supplies a selection argument for why the quiet failure is the one that reaches scale: loud failures are filtered out before adoption; argued by its paper, not measured
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- LLMs Corrupt Your Documents When You Delegate
- Large Language Model Reasoning Failures
- The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs
- LLMs Get Lost In Multi-Turn Conversation
- RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
- Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
- Knowing Is Not Enough: Information Retrievability as a Precondition to Effective LLM Oversight
- LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries
Original note title
document degradation has a model-tier signature — weaker models delete content while frontier models corrupt it making frontier failures harder to detect