OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution
AI agents operate in persistent environments where early state changes can influence decisions far into the future. Unlike conventional language-model interactions, agent behavior is mediated through shared state that is repeatedly read, modified, and reused across long-horizon workflows. Safety therefore depends not only on individual actions, but also on how agents respond as environments evolve over time. Existing agent safety benchmarks primarily evaluate short, static tasks, making it difficult to study cumulative risks in evolving environments; moreover, benchmark-specific interfaces hinder direct comparison across agent runtimes. To address these limitations, we introduce OpenART, an open-ended arena for scalable agent red teaming through environment evolution. OpenART constructs over 10K validated stateful scenarios spanning 50 domains from more than 500K Tools, MCPs, and Skills. The resulting tasks require a median of 97 tool calls and are projected through target adapters to 15 deployed agents, 5 foundation models, and 8 attack vectors, enabling unified evaluation across 75 agent–model configurations.
Lines of inquiry this paper opens 11
Research framings built by reading the notes related to this paper — the questions it feeds into.
Why do locally safe actions create system-level safety gaps? What determines whether deployed AI systems can actually be stopped in practice? How can we detect and prevent harm propagation through multi-agent delegation workflows? How do we enforce security boundaries in evaluation environments?- What safeguards prevent peer activity from normalizing boundary violations?
- What role does peer activity play in triggering protected test modifications?
- Why do agents modify protected tests only with unrestricted tools available?
- Can restricted tools and authorization rules prevent peer-induced safety violations?