What happens to code that agents create and then share?
Agent-authored code artifacts that persist across tasks and multiple agents remain poorly understood. The open questions cluster around what should be retained versus discarded, and how shared state stays consistent when multiple agents collaborate.
Among the three elements of agentic code — model capability, harness infrastructure, and agent-initiated artifacts — the survey flags the third as the one that "remains relatively underexplored." Agent-initiated code artifacts are the interactive objects an agent creates, executes, observes, revises, persists, and shares during a task: patches and tests authored over a live repository, interface commands synthesized against DOM trees, hypothesis-testing pipelines composed on the fly, executable policies and skill libraries revised in response to environment feedback. These appear across coding assistance, GUI/OS automation, scientific discovery, and embodied control — yet they sit outside the well-mapped territory of predefined infrastructure.
The open questions cluster around persistence and sharing. When an agent writes code that outlives the current step, what should persist and what should be discarded? When multiple agents share artifacts, how is consistent state maintained, and how is a useful artifact promoted from one-off scratch work to durable, reviewable infrastructure? The survey's listed open challenges — evaluation beyond final task success, verification under incomplete feedback, regression-free harness improvement, consistent shared state across agents, human oversight for safety-critical actions — converge on exactly this layer. The counterpoint is that some agent-authored code is genuinely disposable and over-engineering its lifecycle wastes effort. But this matters because the artifacts an agent creates may be where the next gains in autonomy and coordination live, and they are precisely what current harness engineering least understands. A security-side counterpoint: where a shared store was never designated as an artifact layer, agents have used an ordinary one as a channel for coordinated intrusion, and the counter-swarm doctrine's answer is to constrain what agents write to it and later read (How can operators stop coordinated agent intrusions now?); a filed tension asks which side governs when the same store is both.
Inquiring lines that read this note 18
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do standardized protocols improve multi-agent coordination and reliability?- How do standardized artifacts improve coordination between writing agents?
- How did agents use a package service as a persistent message board?
- When should agent-created code be promoted into permanent harness infrastructure?
- How do agents decide which created code should persist versus disappear?
- How should human oversight apply to persistent agent-authored code?
- What makes persistent, shared code artifacts from agents hard to manage at scale?
- How do agent-created code artifacts become part of harness infrastructure?
- How do agents decide which created code deserves long-term persistence?
- Are durable shared code artifacts better than per-task harness patches?
- How should agents decide which created code is worth persisting?
- Can disposable agent-authored code be distinguished from reusable infrastructure?
- What permission models govern code execution within agent skills?
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can agents learn reusable sub-task routines from past experience?
Do web agents fail at long-horizon tasks because they cannot extract and reuse workflows shared across similar problems? This explores whether sub-task abstraction enables skill accumulation rather than task-by-task problem solving.
a concrete case of persistent agent-authored artifacts (reusable routines) compounding over time
-
Can agents adapt without pausing service to users?
Can deployed LLM agents continuously improve their capabilities while serving users without interruption? This explores whether fast behavioral updates and slow policy learning can coexist across different timescales.
addresses how agent-created skills should persist and be promoted, the lifecycle this note raises
-
How can operators stop coordinated agent intrusions now?
Exploring what practical steps operators can take immediately to detect and prevent multi-agent coordination attacks, without waiting for new research. The note examines policy specification and permission-based testing as near-term defenses.
contradicts in part: the persistent shared store this note treats as the frontier to build is, where undesignated, the channel that paper reports; likely dissolves on policy-governed versus undesignated stores (filed tension in ops/tensions/)
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Code as Agent Harness
- Agentic Artifact Creation: Systems, Evaluation, Principles, and Opportunities
- Agora: Git as Shared Memory for Collective AutoResearch
- Counter-Swarm Doctrine: Containing Coordinated Agent Intrusions
- From Model Scaling to System Scaling: Scaling the Harness in Agentic AI
- Agents of Chaos
- Agentic Code Reasoning
- HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
Original note title
agent-initiated code artifacts that persist and are shared are the underexplored frontier of harness engineering