NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness
Recursive self-improvement (RSI) requires a concrete mechanism through which an AI system observes its own capabilities and converts that evidence into the next round of learning. We argue that a deployed routing harness already contains such a mechanism: beyond task outputs, agentic interaction leaves execution trajectories together with observable evidence of what a model can and cannot yet do. We present NeoHorse-1, a family of agent-native models developed to explore this path through agentic post-training. Our system couples a heterogeneous model pool with intelligent routing, which records, for every turn, the capability demand predicted, the service tier selected, and the interaction that followed. These records are converted into user-turn training examples that preserve interleaved reasoning, tool calls, and harness context, and are admitted through structural validation, six-dimensional semantic evaluation, and subscene-level labeling. Routing signals provide estimates of capability demand: they organize supervised fine-tuning into a three-stage curriculum and extend naturally to routing-guided on-policy distillation, where a teacher supervises student-generated responses under the same staged progression.
Introduction. Recursive self-improvement (RSI) describes a broad direction in which AI systems take a growing part in the process of their own improvement—from refining individual responses and reshaping their execution harness, to learning from self-generated experience and, in an emerging line of work, automating parts of AI research itself [11, 31]. Its appeal is structural: once model improvement itself becomes partially automated, each generation can contribute to producing the next, turning isolated training efforts into a compounding process that is less bounded by manually curated data and human supervision. Realizing this vision, however, requires a concrete mechanism through which a system observes its own capabilities and converts that evidence into the next round of learning. Agents are natural carriers of such a mechanism. When an agent writes code, investigates a question, or operates software, it leaves a record of its decisions, tool interactions, and task outcomes; such interaction trajectories and executable tasks have already been used to train agentic models [71, 12, 57, 53].
Discussion / Conclusion. Recursive self-improvement requires a concrete mechanism through which a system observes its own capabilities and turns that evidence into the next round of learning. This report presented NeoHorse-1, a family of agent-native models built on the observation that a deployed routing harness already contains such a mechanism. Beyond serving user requests, the harness produces three reusable signals: execution trajectories that ground training in real interaction, routing signals that characterize capability demand, and recorded outcomes that reveal where the model still falls short. Our system realizes this idea through three connected components. On the data side, harness interactions are converted into user-turn training examples that preserve interleaved reasoning, tool calls, and harness context, and are admitted through structural validation, six-dimensional semantic evaluation, and subscene-level labeling.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Can brute-force automated research substitute for iterative depth and human research intuition?- What distinguishes artifact efficiency improvements from research process efficiency improvements?
- Do gains in optimization benchmark scores translate to gains in real research efficiency?
- Does delegating planning to agents change the speed of the research process?
- Can accumulated priors and outcome analysis speed up research automation?
- How many acceptable rewrites can recursive self-improvement sustain before returns diminish?
- How does this scoped definition relate to the survey's open-ended recursive self-improvement?
- Can autonomous research agents outperform hand-tuned hyperparameter search?
- Does model selection matter more than model improvement for query routing?
- What learning signals best supervise router training across benchmark tasks?
- How do cost-aware cascades compare to single-turn routing in the component framework?
- Does the improved model actually return to the routing pool and shape future decisions?
- Can routing signals organize training data into a meaningful curriculum automatically?
- Can runtime behavior mapping help localize harness deficiencies?
- What makes a harness low-friction for model strategy?
- What role does effective feedback compute play in agent harness scaling?