SYNTHESIS NOTE
Topics›Context Engineering›this note

Can an external manager handle context for frozen agents?

Exploring whether a separate trained system can effectively manage a frozen agent's context window. This matters because many deployed agents are closed-source and can't be retrained, yet they suffer from context degradation.

Synthesis note · 2026-06-03 · sourced from Context Engineering
How does test-time scaling work for individual research agents?

Long-horizon agents accumulate context — tool results, intermediate reasoning — until stale content obscures salient evidence, amplifies positional bias, and degrades decisions. Prior fixes put the burden of managing context on the agent itself (agent-side control, or fixed summarization), which requires training the agent and is impractical for closed-source agents, and ignores that different agents need different strategies.

AdaCoM separates the concern entirely: train an external LLM to manage the context of a frozen agent through flexible modification actions and end-to-end RL. The manager prunes stale content while preserving task constraints and progress, improving diverse agents on web-search and deep-research benchmarks and transferring to unseen agents of similar capability.

The most useful finding is a fidelity–reliability trade-off. Agents with higher vanilla ReAct performance benefit from higher-fidelity context preservation — they can use more detail well. Lower-performing agents require more aggressive compression to stay within a reliable reasoning regime. The right amount of context is not a property of the task alone; it is indexed to the agent's own competence. This means context management is not one universal policy but a per-agent calibration — consistent with Does fixed sparsity work for all sequence lengths?, where the optimal budget is also conditional rather than fixed.

Inquiring lines that read this note 44

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How should designers communicate what AI systems truly are and can do? Should agents decouple planning from perception grounding for better performance? How should agents manage memory granularity to improve long-term performance? Can compression size predict model complexity better than parameter count alone? Can memory architectures handle ultra-long context better than attention? How does harness optimization generalize across different model architectures and domains? How do agent-learned skills transfer and improve across different tasks? Why does memory consolidation cause performance regression in continual learning? How should agent systems validate and persist generated code artifacts? When should work require human-AI partnership versus full automation? Can harness architecture and protocols provide agent reliability without model scaling? How should systems decide whether to retrieve or reason alone? Do reasoning benchmarks predict model performance in long-horizon workflows? How do standardized protocols improve multi-agent coordination and reliability? What fundamental constraints limit how effectively agents can improve themselves? What capability trade-offs arise from domain specialization through fine-tuning?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 107 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

context management can be offloaded to a trained external manager for a frozen agent and optimal compression depends on the agent's own reliability