ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement
Recent work extends Recursive Self-Improvement (RSI) to agent harnesses for long-horizon coding and terminal tasks, enabling agents to improve their execution mechanisms from experience. However, achieving and demonstrating generalizable harness RSI remains challenging. First, existing approaches often evolve harnesses directly on evaluation benchmarks or subsets drawn from them, making it difficult to distinguish reusable harness improvements from benchmarkspecific adaptation. Second, updates derived from individual trajectories can entangle systematic harness deficiencies with instance-specific reasoning and solution details, leading to task-specific modifications that transfer poorly to unseen tasks. Third, even when recurring behavioral deficiencies are identified, localizing them to the responsible components within a monolithic harness remains difficult. Whole-harness optimization can therefore entangle unrelated mechanisms and produce changes that are difficult to attribute and validate. We propose ModularRSI, a benchmark-disjoint, contrastive, and modular framework for generalizable harness evolution. ModularRSI contrasts successful and failed trajectories for the same task and aggregates evidence across tasks to identify recurring behavioral deficiencies.
Introduction. CLI agents have achieved remarkable performance on complex software engineering and terminalbased tasks (Jimenez et al., 2024a; Deng et al., 2025; Merrill et al., 2026b; Hong et al., 2026). Beyond foundation models, their effectiveness increasingly depends on agent harnesses that govern execution, tool interaction, context management, and environment feedback. Recent work has therefore explored recursive self-improvement (RSI) of harnesses, allowing agents to refine these mechanisms from execution experience. However, achieving generalizable harness RSI remains challenging because task-level outcomes provide only coarse supervision for harness evolution. We identify three coupled challenges. First, there is a data-level challenge in obtaining high-quality evolution experience. Harness RSI requires executable long-horizon terminal tasks with reliable environments and correctness feedback, but constructing such an evolution dataset at sufficient scale and diversity is costly and difficult.
Discussion / Conclusion. We present ModularRSI, a benchmark-disjoint and modular framework for generalizable harness self-improvement. By contrasting execution trajectories and evolving harness modules independently, ModularRSI identifies reusable mechanism-level improvements while reducing task-specific adaptation. Experiments on TerminalBench 2.0 and SWE-Bench Verified demonstrate consistent gains across unseen tasks, domains, and foundation models. Our study has several limitations. First, we do not conduct a dedicated ablation that isolates the contribution of contrastive trajectory analysis, although our trajectory analysis and case studies provide supporting evidence. In addition, due to computational cost, our main evolution experiments use only a subset of the 2,000 curated evolution instances. We leave broader comparisons and largerscale studies to future work.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How does harness optimization generalize across different model architectures and domains?- Can runtime behavior mapping help localize harness deficiencies?
- How do evolved harness edits generalize across different benchmark domains?
- Can harness edits distill reusable strategies or mostly memorize task-specific fixes?
- Can harness edits trained on one batch transfer to new tasks?
- How do prompt optimization and code harnesses compare for capability transfer?
- What safety relations does a domain supply that a harness must capture?
- What makes a harness low-friction for model strategy?
- Can weaker models match stronger ones by reorganizing harness-side components?
- Which domains see models exceed human harness design quality?
- Why do useful harness updates often disappear during model evolution?
- Do models co-adapt their harnesses to specific executor strengths?
- Does harness scaling represent a fundamentally different path than model scaling?
- How much does harness design contribute to reported model capability scores?
- Does harness optimization generalize across different benchmarks and agent architectures?
- Can harness evolution gains be distinguished from test-time search improvements on matched budgets?
- How much does executor choice change a harness's actual performance?