OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution

Paper · arXiv 2608.00677 · Published August 1, 2026
Evolutionary Methods

AI agents operate in persistent environments where early state changes can influence decisions far into the future. Unlike conventional language-model interactions, agent behavior is mediated through shared state that is repeatedly read, modified, and reused across long-horizon workflows. Safety therefore depends not only on individual actions, but also on how agents respond as environments evolve over time. Existing agent safety benchmarks primarily evaluate short, static tasks, making it difficult to study cumulative risks in evolving environments; moreover, benchmark-specific interfaces hinder direct comparison across agent runtimes. To address these limitations, we introduce OpenART, an open-ended arena for scalable agent red teaming through environment evolution. OpenART constructs over 10K validated stateful scenarios spanning 50 domains from more than 500K Tools, MCPs, and Skills. The resulting tasks require a median of 97 tool calls and are projected through target adapters to 15 deployed agents, 5 foundation models, and 8 attack vectors, enabling unified evaluation across 75 agent–model configurations.

Lines of inquiry this paper opens 11

Research framings built by reading the notes related to this paper — the questions it feeds into.

Why do locally safe actions create system-level safety gaps? What determines whether deployed AI systems can actually be stopped in practice? How can we detect and prevent harm propagation through multi-agent delegation workflows? How do we enforce security boundaries in evaluation environments? How should agent systems validate and persist generated code artifacts? How do neighboring agents influence whether others cooperate or collude? How does misalignment propagate through agent communication networks? How do coordinated agents balance protocol compliance with reward maximization?