Agentic Artifact Creation: Systems, Evaluation, Principles, and Opportunities
Abstract. Generative models can turn natural-language prompts into images, text, code, and other content, lowering the cost of producing drafts and components. Their practical impact increasingly depends on whether those pieces can become complete, dependable deliverables. Such deliverables pose a different challenge because their requirements interact, and failures visible in the final output may be difficult to trace or repair. This survey examines agentic artifact creation, which we define as stateful construction in which an AI system materially constructs or revises a deliverable and intermediate observations redirect later work. Functionally, the process links an operational representation of the artifact, a construction policy, and runtime verification whose feedback can redirect later actions. This control structure can expose dependencies and support targeted revision, but only when observations identify failures at a scope that the available actions can repair. We reviewed 259 works available through August 20, 2026: 230 systems meeting this definition and 29 benchmarks of agentic artifact construction. We compare six artifact families, then analyze application settings and evaluation practice as separate dimensions.
Introduction. Generative models can now produce a wide range of content from natural-language prompts [1, 2]. By lowering the cost of producing drafts and components, this capability is entering a wider range of creative and professional workflows; controlled and field studies already report productivity gains in professional writing and customer support [3, 4]. In these settings, success increasingly depends on complete artifacts whose parts must work together rather than isolated outputs. Here, artifact denotes an intentionally produced, purpose-directed deliverable [5]. Direct generation works well for bounded tasks whose outputs are easy to inspect and inexpensive to regenerate. It is less reliable when the deliverable is governed by several acceptance criteria. Paper2Poster [6], for example, illustrates how a scientific poster must preserve source content, fit a constrained page, and communicate the paper clearly. These requirements call for different kinds of evidence; a plausible image or scalar quality score cannot establish that the poster is ready to deliver.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How do standardized protocols improve multi-agent coordination and reliability?- How do standardized artifacts improve coordination between writing agents?
- How did agents use a package service as a persistent message board?
- When should agent-created code be promoted into permanent harness infrastructure?
- How do agents decide which created code should persist versus disappear?
- How should human oversight apply to persistent agent-authored code?
- What makes persistent, shared code artifacts from agents hard to manage at scale?
- How do agent-created code artifacts become part of harness infrastructure?
- How do agents decide which created code deserves long-term persistence?
- Are durable shared code artifacts better than per-task harness patches?
- How should agents decide which created code is worth persisting?
- Can disposable agent-authored code be distinguished from reusable infrastructure?
- What permission models govern code execution within agent skills?
- Why does embedding research tools in coding assistants improve reliability?
- What prevents multiple agents from corrupting shared state in live artifacts?
- What breaks when multiple agents share and revise the same artifacts?
- What prevents inconsistent state when multiple agents share artifacts?
- How do shared artifact stores become security risks in multi-agent systems?