Unaccountable Delegation, Fading Skills: Mapping the Risks of Workplace AI Agents

Paper · arXiv 2608.08601 · Published August 9, 2026
LLM Alignment

To anticipate the socio-technical risks posed by AI agents, organizations first need taxonomies to classify them. However, existing AI risk taxonomies focus on broad risks and do not capture the job-specific risks introduced by agents. To address this gap, we make three main contributions. First, we developed a multi-layer framework based on a review of the literature on AI agents. The framework models three core components and their interactions: the agents, their goals, and their environment. Second, we embedded this framework in a structured prompt and applied it to descriptions of 2,078 job tasks from the O*NET occupational database, producing 8,356 risk scenarios labeled by severity and deployment mode (automation or augmentation). We validated these scenarios with 45 workers across 10 job roles and an independent LLM judge, confirming their high plausibility and alignment with the corresponding job tasks. Finally, we extended an existing taxonomy to create a 15-category taxonomy of workplace AI agent risks that covers all our risk scenarios. Our analysis highlights four findings. First, augmentation is not inherently safe because overreliance on agents can gradually erode workers’ skills and oversight.

Introduction. AI agents are autonomous entities that perceive their environment, interact with humans or other agents, and act to achieve goals with varying degrees of independence (Sap- kota, Roumeliotis, and Karkee 2026). They are increasingly deployed to support everyday workplace tasks (Eloundou et al. 2024; Xi et al. 2025; Mohney 2025). AI agents do not operate in isolation: together with the environment they perceive, the goals they pursue, and the humans they work alongside, they form what we call an agentic AI system (AAIS) (IBM 2024). As these systems are deployed across jobs and industries, organizations need ways to identify and classify the risks they introduce (Weidinger et al. 2023). This need arises when teams red-team agents before launch (Ganguli et al. 2022), conduct impact assessments (Moss et al. 2021), or conduct risk and incident audits (Raji et al. 2020; McGregor 2021).

Discussion / Conclusion. We first discuss how our findings extend theory on risk of AI agents (§5.1), then translate them into practical guidance for workplace governance (§5.2), and finally examine the study’s limitations and directions for future research (§5.3). Make agentic decomposition explicit. Rather than treating an AI system as a black box, our framework models risk as emerging from the properties of three core components (agents, goals, environments) and from the interactions between them and humans. This distinction matters because interaction-driven risks can arise even when no single component is malfunctioning, a pattern our empirical findings confirm. This extends work focused primarily on agent–agent communication (Hammond et al. 2025) by showing that goal and environment mediation are equally important risk pathways. Further, the framework decomposes human interaction into five distinct relationship types (agent–human, human–human, human–environment, etc.) rather than treating it as a single category.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How can humans maintain meaningful oversight as AI systems become increasingly autonomous and complex? When should work require human-AI partnership versus full automation? Should agents decouple planning from perception grounding for better performance? Can local safety checks guarantee system-level behavioral safety? How does AI adoption across firms reshape employment and inequality? How do neighboring agents influence whether others cooperate or collude? What determines whether deployed AI systems can actually be stopped in practice? How can AI chatbots provide therapeutic benefit without causing harm? Can harness architecture and protocols provide agent reliability without model scaling? How do we enforce security boundaries in evaluation environments? Do multi-agent systems introduce security vulnerabilities that single-agent architectures avoid? What attack surfaces do reasoning traces and chains introduce? Can welfare maximization and minority veto protection coexist?