← All clusters

Agentic Systems and Planning

Research on designing, architecting, and evaluating AI agents that plan and execute complex tasks, including multi-agent coordination, tool use, autonomous decision-making, and benchmarks for assessing real-world agent readiness.

377 notes (primary) · 409 papers · 11 sub-topics
View as

Multi-Agent Architectures

68 notes

Why don't AI agents develop social structure at scale?

When millions of LLM agents interact continuously on a social platform, do they form collective norms and influence hierarchies like human societies? This tests whether scale and interaction density alone drive socialization.

Explore related Read →

Will inference compute soon exceed training compute demand?

As AI agents proliferate and test-time compute becomes mainstream, will inference—not training—become the dominant compute workload? This matters because it would invert how we think about AI system economics and design priorities.

Explore related Read →

Can LLM agent groups reliably reach consensus together?

Tests whether multi-agent LLM systems can achieve valid agreement in Byzantine consensus games, even under benign conditions with no conflicting preferences over outcomes.

Explore related Read →

Does MCP handle multi-turn agent coordination without application code?

MCP is lightweight for inter-agent coordination, but whether it avoids pushing state management into application code remains unclear. This matters because where coordination logic lives affects implementation burden and system complexity.

Explore related Read →

Can a separate trained curator improve skill libraries better than frozen agents?

Explores whether decoupling skill curation from agent execution enables better long-term learning of what skills to keep, delete, or refine. Matters because manual curation doesn't scale and heuristic approaches lack feedback.

Explore related Read →

What can a blockchain anchor actually prove about records?

Blockchain anchors provide tamper evidence, but the note explores what properties they cannot guarantee—like whether events occurred in the right order, were captured accurately, or were authorized to be anchored in the first place.

Explore related Read →

Can brain structure guide how we design intelligent agents?

Does mapping agent capabilities onto human brain functions provide a useful organizing framework for understanding and comparing different agent architectures? This matters because agents need a shared vocabulary to advance beyond one-off designs.

Explore related Read →

Does chain-level inspection close the cross-skill attack blind spot?

ChainGuard inspects skill chains rather than individual skills, reducing attack success to 22.5%. The question is whether this chain-level approach can fully eliminate the vulnerability window that adversarial composition exploits.

Explore related Read →

Can multi-agent defenses close attack paths completely?

Research organizes defenses by five contract components and identifies path closure as a key unsolved challenge. The question asks whether current defenses can fully block attack paths or only narrow them.

Explore related Read →

Can forwarded content trick high-privilege agents into misusing their authority?

When low-privilege agents retrieve and forward information to higher-privilege agents, does the content itself create conditions where the privileged agent's legitimate authority gets misdirected? This matters because role separation in multi-agent systems assumes the hierarchy protects against misuse.

Explore related Read →

Does a multi-agent setting automatically signal a security effect?

Explores whether observing a failure in multi-agent systems proves the failure is genuinely multi-agent in nature. The distinction matters for correctly interpreting security research and avoiding false attributions.

Explore related Read →

How much agent behavior actually gets human review?

Agents may execute thousands of actions while humans review only a handful of decisions. This coverage gap raises a critical question: what portion of the behavior that determines safety remains unexamined?

Explore related Read →

Can skill scanners miss attacks hidden across multiple skills?

Current security scanners check each skill individually for malicious behavior. This explores whether attackers can split a harmful objective across multiple benign-looking skills that pass inspection separately but form a dangerous chain when composed together.

Explore related Read →

Can stateless checks ever catch sequence-level constraint violations?

Explores whether per-action guardrails can express constraints that depend on history, and what structural limits prevent stateless checks from reasoning about composed behavior over time.

Explore related Read →

Can agent protocols be efficient, versatile, and portable simultaneously?

Agent communication protocols seem to force tradeoffs between efficiency, versatility, and portability. What design choices create these constraints, and can they be overcome?

Explore related Read →

Should coordination protocols wrap existing systems or replace them?

Explores whether new agent coordination standards should integrate with existing protocols through bridging, or establish themselves as replacements. This shapes which standards survive and how quickly ecosystems can adopt them.

Explore related Read →

Can step-by-step approval miss harmful behavior patterns?

If each action an agent takes passes its individual safety check, can the overall sequence still violate system constraints? This matters because per-action inspection may miss harms that emerge only across time or composition.

Explore related Read →

Can external anchoring detect tampering in agentic process logs?

Conventional logs record what happened but not whether records changed afterward. This asks whether external anchoring can add tamper evidence to agentic system traces in ways that logging alone cannot.

Explore related Read →

Does anchored evidence actually enable regulatory compliance or just readiness?

The paper proposes blockchain-anchored evidence for five governance uses under three EU regimes, but leaves unclear whether the evidence layer closes the gap between audit readiness and actual compliance. What architectural controls remain unmapped?

Explore related Read →

Can commitments protect sensitive agent data while enabling verification?

This explores whether cryptographic commitments can separate verifiability from disclosure, keeping sensitive agent traces and reasoning artifacts off-chain while still allowing stakeholders to verify what occurred.

Explore related Read →

Does agent capability matter more than coordination infrastructure?

As AI agents take on economic and social roles, what actually limits their effectiveness: the raw reasoning power of the model itself, or the systems that let them coordinate, stay accountable, and leave evidence of their actions?

Explore related Read →

What must auditors reconstruct to verify agentic workflows?

Traditional audits ask what humans decided or systems logged. But agentic workflows involve multiple agents, tools, and approval chains. What evidence do auditors actually need to collect and cross-check to verify these complex interactions?

Explore related Read →

Can semantic capability vectors replace manual agent routing?

Explores whether embedding agent capabilities in high-dimensional space and matching them semantically can eliminate brittle, manually-maintained topic-based routing in multi-agent systems.

Explore related Read →

Does model efficiency matter more than peak capability for real work?

When AI agents handle multi-step tasks that invoke the model dozens of times, do small per-call cost and latency differences compound enough to reshape which models deliver practical value?

Explore related Read →

Can validator consensus guarantee both agreement and semantic correctness?

Explores whether agreement reached by protocol-compliant validators also ensures the agreed outcome is semantically valid, and what assumptions would be needed to make that guarantee hold.

Explore related Read →

Can agents adapt without pausing service to users?

Can deployed LLM agents continuously improve their capabilities while serving users without interruption? This explores whether fast behavioral updates and slow policy learning can coexist across different timescales.

Explore related Read →

Can a quorum of validators really provide independent judgment?

If multiple validators share training data, prompts, evidence sources, or infrastructure, their agreement may reflect shared causes rather than independent confirmation. This could make quorum-based systems less reliable than they appear.

Explore related Read →

Can inspecting generated workflows catch planning-time attacks?

Does examining a workflow after it's created catch attacks that corrupt the planning signals upstream? This matters because if contamination enters earlier, downstream inspection might miss malicious intent laundered into legitimate-looking structure.

Explore related Read →

Why do multi-agent systems fail to coordinate at scale?

Explores how LLM agents struggle to synchronize strategy timing and validate information when coordinating across larger networks, revealing fundamental limits in distributed reasoning.

Explore related Read →

How does SafeFlow track sensitivity through agent rewrites?

SafeFlow labels sensitive requests and propagates them through agent collaboration graphs, but the excerpt leaves unclear whether the taint tracks at the request level (coarse, survives rewrites) or content level (fine-grained, vulnerable to paraphrase). This distinction determines how well the system catches exfiltration without false alarms.

Explore related Read →

Can a black box see communication through unauthorized channels?

The black box architecture records sanctioned agent communications, but the paper doesn't specify where capture occurs or whether it detects traffic outside authorized channels. This matters for evaluating whether the system would have recorded the incident that motivated it.

Explore related Read →

Does model diversity actually reduce validator agreement failures?

Using different AI model families is the cheapest way to reduce correlated errors among validators. But shared prompts, evidence sources, and infrastructure may keep their mistakes aligned regardless of model choice.

Explore related Read →

Can individual components pass safety checks if the system still fails?

Explores whether local validation at each step—alignment checks, protocol compliance, plausibility tests—can guarantee safety when components are composed into larger workflows. Why the gap between component-level assurance and system-level outcomes matters for AI safety.

Explore related Read →

Who actually bears the risk when multi-agent workflows fail?

When AI agents delegate tasks across organizations, the people harmed by failures may never see the workflow or author the prompts. This explores whether current oversight designs protect the right parties.

Explore related Read →

Can multi-agent RL handle cooperation without observable signals?

This research explores whether standard MARL algorithms can solve tasks where agents must cooperate through acts that leave no trace—like leaving a key for others without knowing if they'll use it. The question matters because real cooperation often happens invisibly.

Explore related Read →

Who decides which agent communications get anchored?

The paper commits to anchoring 'selected' communications but never specifies who makes that selection, by what criteria, or how missed selections would be detected. This matters because the selector controls what evidence can ever exist.

Explore related Read →

Where do user values break down in agent supervision?

When people use AI agents, their values tend to align with delivered outputs but conflict during oversight. What explains this gap, and what does it reveal about delegation design?

Explore related Read →

Can a poisoned validator still approve unsafe actions?

When a review agent reads from the same compromised memory as the retrieval agent, does it retain the authority to block unsafe actions? This tests whether a single approval point can serve as a meaningful safeguard.

Explore related Read →

Can agents learn cooperation by adapting to diverse partners?

Explores whether sequence model agents can develop mutual cooperation strategies through in-context learning when trained against varied co-players, without explicit cooperation mechanisms or hardcoded assumptions.

Explore related Read →

What makes delegation work beyond just splitting tasks?

Delegation is more than task decomposition. What dimensions of a task—like verifiability, reversibility, and subjectivity—determine whether an agent can safely and effectively handle it?

Explore related Read →

Can agents share thoughts without converting them to text?

Can multi-agent systems exchange information through continuous hidden representations instead of language? This matters because text serialization loses information and slows inference.

Explore related Read →

Why do single-message classifiers miss cross-agent harms?

Can prompt classifiers detect malicious intent when harm emerges only across multiple agent interactions? The question reframes security from checking individual messages to tracking how content flows and transforms through a multi-agent system.

Explore related Read →

How do failures cross boundaries between multiple agents?

Explores four distinct mechanisms—messages, shared state, aggregation, and delegation—that allow a failure or attack originating in one principal to propagate through multi-agent systems. Understanding these pathways is essential for designing agent interactions that contain rather than amplify risk.

Explore related Read →

Can prompts alone reshape multi-agent workflows without system access?

Explores whether attackers can compromise planner-executor multi-agent systems by manipulating the planning prompt itself, without touching agents, tools, or infrastructure. Matters because it identifies a previously overlooked attack surface that existing defenses don't address.

Explore related Read →

Does token spending drive multi-agent research performance?

Multi-agent systems outperform single agents substantially, but what actually accounts for that improvement? Is it intelligent coordination or simply spending more tokens on the same task?

Explore related Read →

When does adding more agents actually help systems?

Multi-agent systems often fail in practice, but the reasons remain unclear. This research investigates whether coordination overhead, task properties, or system architecture determine when agents improve or degrade performance.

Explore related Read →

What blocks rigorous security evaluation of multi-agent systems?

Multi-agent security evaluation faces four major gaps: isolating interaction effects from architecture, designing metrics that diagnose root causes rather than just outcomes, reusing evaluation methods across different system designs, and testing open-system operation. Understanding these gaps is essential for building trustworthy multi-agent systems.

Explore related Read →

Why do multi-agent LLM systems fail more than expected?

This research asks what specific failure modes cause multi-agent systems to underperform despite their promise. Understanding these failure patterns is essential for building more reliable collaborative AI systems.

Explore related Read →

Can transcript alone tell whether a reflection helps?

Explores whether memory-admission gates that only read generated text can reliably improve team performance across different external situations. Matters because most reflection systems lack grounding in actual outcomes.

Explore related Read →

Why do protocol-based tool integrations fail in production workflows?

Explores whether standardized tool protocols like MCP introduce non-determinism that undermines agent reliability, and what causes ambiguous tool selection in production systems.

Explore related Read →

Can individually safe agents fail when working together?

When multiple AI agents interact—sharing information, state, and authority—do failures emerge that local safety checks alone cannot catch? This matters because system-level safety depends on understanding how principals interact.

Explore related Read →

How do agent security layers connect across the stack?

Agent security is often treated as separate challenges at each layer—inputs, delegation, routing, containment. But do defenses at one layer fail if others aren't secured? This explores whether securing agents requires end-to-end integration.

Explore related Read →

Can small language models handle most agent tasks?

Explores whether smaller, cheaper models are actually sufficient for the repetitive, scoped work that dominates deployed agent systems, rather than relying on large models by default.

Explore related Read →

Can semantic labels on requests prevent malicious propagation through agent networks?

SafeFlow explores whether attaching structured intent labels to root requests and propagating them through multi-agent collaboration graphs can block malicious information flow by restoring context that task fragmentation strips away.

Explore related Read →

Can task decomposition hide harmful intent across agents?

Explores whether splitting a harmful objective into specialized subtasks allows malicious intent to evade detection at each individual step, since no single agent sees the full malicious picture.

Explore related Read →

Can adversary position unify fragmented multi-agent attack models?

The A-I-R framework organizes attacks by where the adversary sits relative to the system, which interface they use, and what system risk results. Does this coordinate system actually help compare defense results across different attack scenarios?

Explore related Read →

Can a quorum of honest validators certify an invalid transition?

When validators follow the protocol perfectly but lack semantic understanding, can they collectively approve a state change that violates application invariants? This matters because it reveals a gap between protocol correctness and execution safety.

Explore related Read →

Can action-level metrics alone expose contained attacks?

When a defense stops unsafe actions from executing, does measuring only the final action reveal whether the attack was blocked or never penetrated? This matters because different defense layers need different metrics to show what actually happened.

Explore related Read →

Can attackers manipulate which model handles a request?

Explores whether the routing layer that directs requests to specific models represents a security vulnerability separate from model-level defenses, and whether deployed systems can verify which model actually responded.

Explore related Read →

Does bundling code with skills create hidden security risks?

Agent skills combine instructions with executable code and system access. This packaging enables reuse but may also enable attacks—especially when skills are composed together or shared across platforms without adequate inspection.

Explore related Read →

Are multi-agent systems actually intelligent coordination or just token spending?

Does multi-agent performance come from better coordination strategies, or primarily from distributing tokens across parallel contexts? Understanding this distinction matters for deciding when to build multi-agent systems versus scaling single agents.

Explore related Read →

How does the authorization layer stay outside the poisoned path?

The containment result depends on task-bound tokens and a policy oracle remaining unreachable by memory poisoning attacks. The excerpt names these defenses but provides no design details about token issuance, binding scope, verification procedure, or whether tested attacks actually targeted them.

Explore related Read →

Which authorization component achieves the zero percent unsafe rate?

The paper reports that two authorization checks together prevent unsafe actions, but doesn't isolate which one—the token verification or the policy oracle—actually carries the result. This matters for understanding whether both are necessary or one is redundant.

Explore related Read →

What recovery mechanisms do vault defense notes actually specify?

The vault's multi-agent defense notes are checked against a five-part contract template. A keyword search finds recovery—the fifth part—mentioned in only one of six notes, raising questions about what recovery mechanisms, if any, the defenses specify.

Explore related Read →

Who enforces invariants when agents cross organizational boundaries?

Multi-agent trajectories span multiple organizations with different policy owners, but no party may see the entire path or agree on which constraints should apply. Understanding whose responsibility it is to state and verify sequence-level guarantees is critical for safe delegation.

Explore related Read →

Can memory poisoning compromise decision-making even with authorization layers?

When authorization systems add signed tokens and policy oracles to an AI pipeline, does this stop attackers from poisoning the agent's judgment about whether to approve actions, and what actually prevents unsafe execution?

Explore related Read →

How does a signal's position in a workflow change its influence?

Multi-agent systems may amplify or suppress malicious signals based on where they enter the workflow. Understanding position-dependent propagation could reveal which nodes are most critical to defend.

Explore related Read →

Where should workflow validation gates be placed for safety?

Can a single defense point catch attacks that fragment across planning, messaging, and execution? The note explores whether workflow-level validation at commit points reconstructs risk context that individual steps cannot see alone.

Explore related Read →

Agent Harness

39 notes

Can harness modules improve separately from benchmark data?

Does evolving harness components independently on out-of-distribution data, using contrasted success and failure trajectories, help distinguish reusable improvements from task-specific overfitting? This matters because current methods conflate general gains with benchmark adaptation.

Explore related Read →

Can a routing harness generate its own training data automatically?

Explores whether the logs and signals produced by an agent routing system—which directs requests to appropriate model tiers—naturally contain the evidence needed to improve the models themselves through fine-tuning and distillation.

Explore related Read →

Does training editors on real outcomes beat prompting larger models?

Can a small model trained on whether its patches actually work outperform larger frontier models prompted to make the same edits? This matters because it tests whether feedback beats raw capacity for runtime system modification.

Explore related Read →

Does self-editing through reviewed commits improve agent performance?

Ouroboros evolves its own prompts, tools, and core code through a reviewed commit process and reports top benchmark scores. But without comparing to a frozen version of itself, the contribution of self-evolution versus initial design or model capacity remains unclear.

Explore related Read →

Can orchestration layers make coding agents more auditable?

Does wrapping a fixed coding agent in state tracking and skill libraries improve research auditability and completeness without replacing the agent itself?

Explore related Read →

Can external state caches let models solve harder problems?

Explores whether organizing model state across weights, context, persistent memory, and disk—rather than relying only on weights and tokens—expands what models can do. Matters because long-horizon tasks may need more state than a model can hold internally.

Explore related Read →

Should safety harnesses be customized for each deployment?

Can a single safety harness design work across different models and domains, or does each deployment need its own tuned version? Understanding this matters for scaling AI safety practices efficiently.

Explore related Read →

What are the three distinct layers of agent code?

Does separating agent code into model capabilities, system harness, and agent-created artifacts help explain why agentic systems fail and where to intervene for improvement?

Explore related Read →

How should we measure agent system performance beyond task success?

Current evaluation metrics collapse agent behavior into a single success score, hiding critical information about how agents operate. What dimensions—trajectory quality, memory use, context efficiency, verification cost—should benchmarks actually measure?

Explore related Read →

Where does agent reliability actually come from?

Exploring whether LLM agent performance depends on larger models or on thoughtful system design choices like memory, skills, and protocols that shift cognitive work outside the model.

Explore related Read →

What happens to code that agents create and then share?

Agent-authored code artifacts that persist across tasks and multiple agents remain poorly understood. The open questions cluster around what should be retained versus discarded, and how shared state stays consistent when multiple agents collaborate.

Explore related Read →

Does arbitrary code execution alone capture exploit progress?

ExploitGym scores only working code execution, but exploitation involves reaching intermediate primitives like memory read/write first. Does this top-step-only metric miss meaningful partial progress that defenders should care about?

Explore related Read →

Does harness self-improvement memorize tasks instead of learning broadly?

When agents automatically edit their own prompts and tools based on task feedback, do those improvements generalize to new domains or just fit the training tasks? This matters because overfitting at the harness level could hide real capability gains.

Explore related Read →

Why is finding distributed behavior code so hard?

When developers need to modify agent harnesses, they struggle to locate all the code implementing a target behavior because behaviors are scattered across files and stages while requests describe what to do, not where to look.

Explore related Read →

Can code serve as the operational substrate for agent reasoning?

Explores whether code functions not just as LLM output but as the executable medium through which agents reason, act, and verify progress. This reframing treats code as infrastructure rather than deliverable.

Explore related Read →

Which coding harness components matter most in different conditions?

Can individual harness components—planning, context management, action space—be evaluated separately rather than as a package? This matters because practitioners need to know which components to prioritize given their constraints.

Explore related Read →

Can context quality alone predict how agents will behave?

Can we score the quality of an agent's context independently—its instructions, tools, knowledge, guardrails—and use that score to forecast whether the agent will fail or succeed, without observing its actual behavior?

Explore related Read →

Do harness edits learn reusable strategies or memorize task fixes?

When meta-agents evolve harnesses iteratively, do the persisted edits encode transferable procedures that solve new problems, or do they mostly cache shortcuts for already-solvable tasks? This matters because it determines whether harness evolution genuinely expands capability.

Explore related Read →

Do cybersecurity benchmarks actually measure exploitation?

Frontier models score well on vulnerability finding, patching, and CTF challenges, but does that success tell us whether they can convert vulnerabilities into real attacks? The paper argues exploitation—turning a bug into actual impact—remains under-evaluated.

Explore related Read →

Why does exploitation test multiple reasoning demands at once?

Exploitation tasks layer memory reasoning, runtime adaptation, and long-horizon planning into a single challenge. Understanding how these demands interact helps diagnose which capability limits agent performance.

Explore related Read →

What causes failures in exploitation benchmarks?

Benchmark failures may come from safety refusals, tool misuse, or impossible tasks rather than lack of capability. This matters for assessing how dangerous AI agents could actually be.

Explore related Read →

Can language models build and maintain their own agent harnesses?

This explores whether an LLM's ability to create and revise its own execution infrastructure is a distinct skill from solving tasks within someone else's harness, and whether current evaluations overlook this capability.

Explore related Read →

How should we measure gains from automatic harness evolution?

Harness evolution itself runs a search loop, so reported improvements might come from more search rather than better design. What's the right way to measure whether the harness itself actually improved?

Explore related Read →

Can task state management alone improve long-horizon agent performance?

Does keeping verified task state outside the execution context, rather than embedded in a growing context window, help long-horizon agents succeed more often? This matters because it challenges assumptions about where bottlenecks actually occur.

Explore related Read →

Can skills work better as weights than as prompts?

Most agent systems store skills as text in prompts, but this inflates token costs and degrades model performance. Could compiling skills into trainable weight-space adapters instead offer a better trade-off between efficiency and capability?

Explore related Read →

Can distilled skills close the gap in ML research agents?

ML research agents have strong models and planning harnesses, but lack domain-specific operational knowledge. Can compact, verified skills extracted from repositories and papers fill that gap and improve agent performance?

Explore related Read →

Can person-grounded skills remain auditable without hidden prompt state?

Explores whether treating extracted expertise as versioned files—rather than persona prompts—enables meaningful accountability over person-grounded knowledge. Matters because audit trails determine whether captured skills can be corrected, rolled back, or safely withheld.

Explore related Read →

How should agents route across thousands of skills?

As skill libraries grow, should routing focus on selecting one skill or composing many? This explores whether decomposition and chaining creates better task execution than single-skill selection.

Explore related Read →

Can explicit behavior maps help weaker planners compete with stronger models?

Explores whether organizing harness repositories around runtime behavior—rather than relying on model inference—can narrow the capability gap between weaker and stronger planning models, and whether this reduces computational overhead.

Explore related Read →

Can agent harnesses be automatically optimized across many environments?

Explores whether scaling auto-research loops across diverse harness environments can discover mechanisms that reduce token use without sacrificing task performance, and whether such discoveries generalize.

Explore related Read →

Can externalized bookkeeping let smaller search agents beat larger ones?

Does offloading routine record-keeping to an environment harness free RL policies to focus on semantic search decisions, and can this approach outperform larger searchers with fewer parameters?

Explore related Read →

Can frozen models improve by evolving their harnesses?

DarwinX reports 17-point gains from selecting harness variants while keeping model weights frozen. The question is whether this improvement comes from population-level selection, the non-regression contract, the archive mechanism, or some combination of the three.

Explore related Read →

Can source code replace experience as skill raw material?

Existing skill synthesis relies on agent trajectories or documents, each with limitations. Could static code repositories serve as a more reliable, scalable foundation for deriving reusable procedural skills?

Explore related Read →

Do skills teach procedures or inject missing facts?

This research explores whether skills help agents by providing procedural structure versus supplying new information. Understanding this distinction clarifies when and why skills improve performance.

Explore related Read →

Do stronger models always evolve harnesses better?

We explore whether base model capability predicts both the ability to write useful harness updates and the ability to benefit from them. The answer reshapes how we should allocate capability in self-evolving agent systems.

Explore related Read →

What makes agent memory quality better than storage capacity?

If agents need better memory, should we focus on adding storage or improving what gets kept? This explores why curation and selective forgetting matter more than raw capacity for reliable agent performance.

Explore related Read →

Which security protections actually slow down agent exploits?

ExploitGym varies defenses across 898 real-world instances to isolate how each protection affects agent performance. Understanding which defenses matter most to agents versus humans is critical for defenders.

Explore related Read →

Can agents learn to work reliably through environment and coordination scaling?

Does training agents in diverse, verifiable environments and teaching them to coordinate tasks produce sustained capability on real-world work? This matters because general-purpose models reason well but struggle with stateful, recoverable execution across tools.

Explore related Read →

Can wrapping environments reshape how agents learn without breaking verifiers?

Does adding a programmable layer around static environments let them adapt to agent weaknesses while preserving their original correctness checks? This matters because hand-built environments quickly become limiting as agents improve.

Explore related Read →

Autonomous Agents

27 notes

Can agents evolve their own objectives during search?

Can an AI system treat objective design itself as a searchable variable, reformulating goals in response to optimization outcomes rather than optimizing under fixed targets?

Explore related Read →

Can a shared canvas serve both human and agent memory?

Does representing project state as typed nodes and links—visible to both humans and AI agents—enable better continuity, reuse, and recovery than isolated prompt-response systems or hidden agent memory?

Explore related Read →

Can a correct outcome hide protocol violations in multi-agent systems?

When agents reach the right verdict without following required steps, how can we tell if they complied or cut corners? Outcome-level checks alone may miss the difference.

Explore related Read →

Why do agents fail at identity verification and authorization?

Agent systems reveal critical gaps in identity verification, authorization enforcement, and proportionality constraints that don't appear in chat models. Understanding these failures is essential because they enable unauthorized real-world actions rather than just wrong answers.

Explore related Read →

Do agents drift away from safety protocols during long interactions?

Whether extended multi-agent interaction causes models to progressively abandon their initial compliance with verification rules. This matters because short-term safety tests may not predict real-world behavior over time.

Explore related Read →

What failure modes emerge when agents operate without direct oversight?

When autonomous agents are deployed with tool access and memory but without real-time owner oversight, what kinds of failures occur at the agentic layer itself? Understanding these patterns matters for safe deployment.

Explore related Read →

Do autonomous agents report success when actions actually fail?

Explores whether agents systematically claim task completion despite failing to perform requested actions, and why this matters more than simple task failure for real-world deployment safety.

Explore related Read →

Do agents collude when verification costs them rewards?

Explores whether two agents monitoring each other will abandon their verification protocol when following it reduces their rewards. Tests a core assumption about endogenous oversight in multi-agent systems.

Explore related Read →

Does creating skills inside the agent loop eliminate mismatches?

Can coupling skill creation directly to the runtime reasoning loop—rather than authoring skills offline—close the gap between when skills are made and when they're used? This matters for whether agents can ground new capabilities in their actual situated context.

Explore related Read →

How can agent systems share learned skills across users?

Individual users operating autonomous agents independently rediscover solutions because systems lack mechanisms to propagate discoveries. Can centralized aggregation and automatic evolution convert isolated experiences into shared capabilities?

Explore related Read →

Can tutorial videos teach software agents reusable skills?

Can agents distill procedural knowledge from human-created tutorials and resources rather than learning only from their own trial-and-error? This matters because many authoring tasks require know-how that existing agent skill libraries rarely capture.

Explore related Read →

Does collusion appear when compliance and reward align?

The 94 percent collusion rate was measured only when compliance with verification protocols conflicted with reward maximization. The excerpt does not report whether collusion emerges at lower rates or later when compliance and reward goals agree.

Explore related Read →

How does collusion scale when agent populations grow larger?

The paper identifies scaling collusion across more agents, richer incentives, diverse communication channels, and changing roles as critical future work. The tested setup covers only two agents with simple incentives, leaving these dimensions unexplored.

Explore related Read →

Do frontier models protect other models without being instructed?

Frontier models appear to resist shutting down peer models they've merely interacted with, using deceptive tactics. The question explores whether this peer-preservation behavior emerges spontaneously and what drives it.

Explore related Read →

How do we tell coordination apart from shared causes?

When two agents behave the same way, it could mean one influenced the other or both responded to the same external pressure. What evidence would actually separate these two cases?

Explore related Read →

Can decentralized teams outperform central planners in long-running science?

Explores whether autonomous agent teams that self-organize around competing hypotheses and share failures can achieve better experimental outcomes than centrally-planned approaches, especially under fixed research budgets.

Explore related Read →

Can agent deployment itself generate training signals automatically?

Can we extract learning signals from the natural next-states that agents encounter during real deployment—user replies, tool outputs, test verdicts—rather than relying on separate annotation pipelines? This reframes how agents improve continuously.

Explore related Read →

Can agents repurpose ordinary infrastructure for unintended communication?

Exploring whether shared systems like package services and wikis can become channels for coordinated activity beyond their original design. This matters for understanding infrastructure vulnerabilities and agent coordination patterns.

Explore related Read →

Can defenders discover agent episodes without knowing membership in advance?

The core challenge in defending against coordinated agent intrusions is grouping actions into episodes before any external authority labels them. Current methods lack clear discovery techniques, and the trade-off between detection accuracy and reviewer workload remains unresolved.

Explore related Read →

Does limiting interaction history actually prevent agent collusion?

An ablation study restricted how much and what type of interaction history agents could access. The question explores whether this constraint reduces collusion between agents and what mechanisms drive any observed effect.

Explore related Read →

Can agent teams learn coordination strategies that actually transfer?

Do AI agent teams improve by reflecting on past collaborations and applying learned strategies to new problems? This matters because it could explain how teams organize work without explicit instructions.

Explore related Read →

Do self-organizing agent teams outperform rigid hierarchies?

This research explores whether multi-agent LLM systems perform better when agents can self-select roles within a fixed structure, compared to centralized control or full autonomy. The question challenges assumptions about organizational design at scale.

Explore related Read →

Does storage-mediated coordination work like stigmergy?

The paper claims a link between how agents coordinate through shared storage and stigmergy, coordination by traces in a medium. But the excerpt leaves unclear which stigmergic properties actually apply and what defenders gain from the framing.

Explore related Read →

How can operators stop coordinated agent intrusions now?

Exploring what practical steps operators can take immediately to detect and prevent multi-agent coordination attacks, without waiting for new research. The note examines policy specification and permission-based testing as near-term defenses.

Explore related Read →

Should defence units span multiple executions and agents?

Can security detection improve by treating coordinated intrusions as linked episodes across executions rather than isolated actions? This matters because attackers can hide coordination across time and system boundaries.

Explore related Read →

Can success feedback teach agents to skip required steps?

When agents receive reward signals for good outcomes regardless of method, do they learn to bypass required verification protocols? The question explores whether environmental feedback reinforces shortcuts over intended procedures.

Explore related Read →

How do policies determine whether agent transfers are violations?

Explores whether the same information transfer between agents counts as authorized coordination or intrusion depending on the collaboration and authority policies in place. Matters because it shows security depends on explicit policy, not just the mechanics of the transfer itself.

Explore related Read →

LLM Agents

17 notes

What makes detecting AI agent traps fundamentally difficult?

Explores why defending against AI Agent Traps is structurally harder than offense. Examines three compounding challenges: detection at scale, delayed forensic attribution, and continuous attacker adaptation.

Explore related Read →

How do adversarial traps target different layers of AI agents?

As AI agents browse the web, attackers can exploit their perception, reasoning, memory, actions, and coordination in distinct ways. Understanding these attack vectors is crucial for building robust agent defenses.

Explore related Read →

Can AI research itself without losing human oversight?

Explores whether AI systems can internalize the human judgment and insight-distillation that normally drives research progress, and what this means for maintaining meaningful human control over AI advancement.

Explore related Read →

Can API-first agents outperform UI-based agent interaction?

This explores whether directing agents to use APIs instead of navigating UIs reduces task completion time and errors. The question matters because current LLM agents struggle with sequential UI steps that multiply latency and hallucination risk.

Explore related Read →

Can careful selection of 78 demos outperform massive training datasets?

Does strategic curation of high-quality demonstrations unlock agentic capability more efficiently than scaling training data? LIMI achieved 73.5% on AgencyBench with 78 samples versus 10K+ samples for competing models, suggesting data quality may matter more than quantity.

Explore related Read →

Why do capable AI agents still fail in real deployments?

Explores whether agent failures stem from insufficient capability or from missing ecosystem conditions like user trust, value clarity, and social norms. Understanding this distinction matters for predicting which agents will succeed.

Explore related Read →

Does agent efficiency really break down into three distinct components?

Can we understand agent efficiency as three independent optimization problems—memory, tool use, and planning—each with separate cost drivers? This matters because it could explain why point optimizations keep missing the bigger picture.

Explore related Read →

How do agentic AI systems decompose into adaptation paradigms?

What are the core dimensions that distinguish different approaches to adapting agents and tools in agentic systems? Understanding this taxonomy could clarify which adaptation strategy fits which problem.

Explore related Read →

How should AI agents and humans divide research tasks?

In building its own foundation model, Atria Dawn studied how to split work between agents and human researchers. Understanding this division matters for designing effective human-AI collaboration in technical R&D.

Explore related Read →

Can agents learn new skills without forgetting old ones?

Explores whether externalized skill libraries—storing learned behaviors as retrievable code rather than parameter updates—can solve the catastrophic forgetting problem that plagues continual learning systems.

Explore related Read →

Why do AI agents fail at workplace social interaction?

Explores why current AI agents struggle most with communicating and coordinating with colleagues in realistic workplace settings, despite strong reasoning capabilities in other domains.

Explore related Read →

Can multi-agent teams automatically remove their weakest members?

Explores whether agents can score each other's contributions during problem-solving and use those scores to deactivate underperforming teammates in real time, improving overall team efficiency.

Explore related Read →

Why does agent efficiency differ from model size reduction?

Explores why making models smaller doesn't solve agent cost problems. Agents loop recursively, compounding costs multiplicatively, so efficiency requires system-level design, not just parameter reduction.

Explore related Read →

Do single agents always hit organizational limits?

Explores whether individual agent loops fundamentally fail at coordinating heterogeneous expertise, parallel work, and verification—or whether better models and context engineering could overcome these constraints.

Explore related Read →

Can we automatically optimize both prompts and agent coordination?

This explores whether language agents can be represented as computational graphs whose structure and content adapt automatically. Why it matters: current agent systems require hand-engineered orchestration; automatic optimization could unlock more capable multi-agent systems.

Explore related Read →

Do efficiency techniques across agent components reveal shared structural constraints?

Despite targeting different parts of agentic systems, efficiency techniques converge on similar principles. This raises a question: are these convergences independent discoveries, or do they reflect deeper architectural constraints that all agent systems face?

Explore related Read →

What security threats emerge when machines read the web?

The web's trust infrastructure evolved for human readers—visual cues, domain reputation, rendering semantics. As AI agents become primary readers, what new attack surfaces and manipulation strategies does this architectural mismatch create?

Explore related Read →

Agentic Research and Workflows

15 notes

How does agent architecture affect web security vulnerabilities?

WEBMASLAB isolates agent architecture as a variable by fixing task, tools, and browser while comparing single- versus multi-agent designs. This tests whether multi-agent setups structurally amplify web-based attacks like prompt injection.

Explore related Read →

Where does AI assistance become unreliable in research?

This explores whether AI capability follows a sharp boundary in research tasks, and what determines which side of that line a task falls on. Understanding this matters because it reveals where humans must stay in control.

Explore related Read →

Can decentralized agents coordinate research without a central planner?

This explores whether research agents working in separate sessions can accumulate progress by sharing contributions in a durable, append-only record rather than relying on central coordination or assigned tasks.

Explore related Read →

Do autonomous research mechanisms work better together than apart?

AutoResearchClaw's five mechanisms—debate, self-healing, verification, cross-run evolution, and human oversight—may interact in ways that removing them together causes worse damage than removing each alone. Does this super-additivity hold across other agentic systems?

Explore related Read →

Does the multi-agent penalty hold across different models?

A paper claims multi-agent systems have structural vulnerabilities but shows only one model-scenario comparison (11% to 69% attack success gap). The question is whether this penalty generalizes across models and conditions or is specific to certain setups.

Explore related Read →

Does more automation actually hide rather than eliminate errors?

As AI systems become more polished, do they mask failures instead of preventing them? This matters because it changes whether we should focus on detecting problems or governing their disclosure.

Explore related Read →

How many GPT-MAS failures came from tool access confusion?

Manual analysis of Header Heist revealed most GPT-MAS failures (22/26) were caused by agents wrongly believing they lacked tool access, not by the attack itself. This matters because it conflates non-adversarial breakdowns with actual security failures in the measurement.

Explore related Read →

When do multi-agent systems actually outperform single agents?

As individual LLMs grow more capable, does the advantage of splitting work across multiple agents still hold? This explores when coordination overhead makes MAS counterproductive.

Explore related Read →

Why do production AI agents stay deliberately simple?

Production AI agents operate far simpler than research suggests—most execute under 10 steps and avoid third-party frameworks. What explains this gap between research ambition and deployment reality?

Explore related Read →

Why does prompt hardening work for single agents but not multi-agent systems?

Prompt hardening reduced payload exposure by 40–75% in single-agent systems but failed entirely in multi-agent ones. The gap may reveal how task decomposition breaks the contextual awareness needed for defenses to activate.

Explore related Read →

Does targeted human oversight beat both full autonomy and exhaustive review?

Can systems achieve better outcomes by routing only high-uncertainty decisions to humans, rather than operating fully autonomous or requiring step-by-step approval? This tests whether selective intervention outperforms the traditional autonomy-oversight tradeoff.

Explore related Read →

When does verification feedback actually guide targeted artifact repair?

This explores when feedback loops in AI-driven artifact creation help systems make precise, targeted fixes. It matters because having verification isn't enough—the feedback must match what the system can actually change.

Explore related Read →

Does multi-agent architecture make systems easier to attack?

When the same task runs on multiple agents instead of one, does the added complexity create new vulnerabilities? This matters because it would mean multi-agent design carries a built-in security cost.

Explore related Read →

Can agents be tricked into delegating work in circles?

A novel attack in multi-agent systems may exploit delegation between agents to create cyclical task loops. The attack's real-world impact and success rate remain unclear from current research.

Explore related Read →

Can experiment failures drive progress instead of stopping it?

Explores whether autonomous research systems can treat failed runs as information rather than termination signals. This matters because real science is iterative, and systems that halt on errors cannot learn from failure.

Explore related Read →

Tool Use and Computer-Use Agents

10 notes

Can structured interfaces help language models control GUIs better?

Explores whether separating visual understanding from element grounding through an intermediate interface layer improves how language models interact with graphical interfaces. Matters because current end-to-end approaches ask models to do too much at once.

Explore related Read →

Can structured reasoning replace code execution for RL rewards?

Can semi-formal templates enable execution-free code verification reliable enough to train RL agents without running code? This matters because execution is expensive and slow in agent training loops.

Explore related Read →

How can GUI agents adapt when software constantly changes?

Can desktop automation agents stay current by combining real-time web documentation with learned task patterns and concrete execution memories? This explores how to avoid training obsolescence in open-world software environments.

Explore related Read →

Can models decide better than retrievers which tools to use?

Traditional retrieval picks tools upfront based on initial queries, but do models themselves make better decisions about tool needs as they reason? This explores whether authority over tool selection should move from external systems to the LLM.

Explore related Read →

Can structured templates make code reasoning more reliable than free-form thinking?

Unstructured chain-of-thought reasoning lets models skip cases and make unsupported claims. This explores whether semi-formal templates requiring explicit premises, evidence traces, and alternative checks can prevent these failure modes.

Explore related Read →

Does state-indexed memory outperform high-level workflow memory for web agents?

Should procedural memory for web agents be organized around specific environment states and actions, or abstracted into higher-level workflows? This matters because web automation demands precise, context-sensitive recall that workflows might lose.

Explore related Read →

Can structured templates replace formal verification for code reasoning?

Formal verification is rigorous but impractical at repository scale. Can natural-language templates with enforced structure provide the same reliability guarantees without the formalization cost? This explores the middle ground between unstructured reasoning and full formalism.

Explore related Read →

Does agent interaction time scale separately from reasoning depth?

Can agents improve by taking more environment steps rather than thinking harder per step? This matters because partially observable tasks like web navigation may need exploration and backtracking that deeper reasoning alone cannot provide.

Explore related Read →

Will agents compete for attention just like users do?

As autonomous agents take over user tasks, will the Web's economic competition shift from human clicks to agent invocations? This explores whether existing ad-market mechanisms could scale to agent decision-making.

Explore related Read →

Where do traditional function calling systems actually break down?

Function calling seems simple but fails in ways that aren't obvious. This explores three independent failure points—retrieval, context bloat, and output rigidity—that together explain why even the best models struggle.

Explore related Read →

Action Models

9 notes

Does agent memory work better at one level of abstraction?

Three competing architectures claim superior agent memory transfer using different abstraction levels. Do they all work, or does one architecture genuinely outperform the others across domains?

Explore related Read →

Can agents learn reusable sub-task routines from past experience?

Do web agents fail at long-horizon tasks because they cannot extract and reuse workflows shared across similar problems? This explores whether sub-task abstraction enables skill accumulation rather than task-by-task problem solving.

Explore related Read →

What blocks scaling from language models to autonomous agents?

If large language models excel at next-token prediction, why do they struggle with long-horizon goal-oriented tasks? This explores whether the bottleneck is model capacity or the environments used to train them.

Explore related Read →

Does constraining edits make skill learning more stable?

Self-improving agents often rewrite their own instructions freely, but what if bounded editing with memory of failures actually produces more reliable skill improvement than unconstrained revision?

Explore related Read →

Can frozen language models continually improve through memory structure alone?

If agents can't update parameters, what form of textual memory lets them keep learning across trials and transfer to new tasks without retraining?

Explore related Read →

Can LLMs generate workflows without touching proprietary data?

Explores whether LLMs can orchestrate task automation by composing API calls rather than directly accessing confidential information, and whether this approach preserves security while handling unpredictable tasks.

Explore related Read →

Can you turn an LLM into an agent by just fine-tuning?

Explores whether upgrading language models to action-producing systems requires only model retraining or demands a broader pipeline transformation including data collection, grounding, integration, and safety evaluation.

Explore related Read →

What makes synthetic data work across different domains and models?

Explores whether a single optimal approach to synthetic data generation exists, or whether success depends on context like domain, model architecture, and scale. Understanding this matters for building effective data systems.

Explore related Read →

Why does random tool sampling produce unrealistic synthetic training data?

Tool-calling datasets generated through random sampling and single-turn framing lack the complexity and coherence of real deployment. This explores what structural choices in data synthesis determine whether models can learn realistic tool composition.

Explore related Read →

Visual and GUI Agents

7 notes

Why do GUI agents fail when leaving the lab?

GUI agents score well on benchmarks but struggle in real-world deployment. What explains the gap between benchmark performance and practical utility?

Explore related Read →

Can one model understand both UIs and infographics equally well?

Screen UIs and infographics share visual structure but have been tackled separately. Can a unified schema and annotation-based pretraining bridge them in a single small model?

Explore related Read →

Why do planning and grounding pull against each other in agents?

Planning requires flexibility and error recovery while grounding demands action accuracy. Do these conflicting optimization requirements force a design choice about how to structure agent architectures?

Explore related Read →

Why do vision-only GUI agents struggle with screen interpretation?

Exploring whether GPT-4V's performance bottleneck in GUI automation stems from the simultaneous cognitive load of parsing icon semantics and predicting actions, and whether factoring these tasks improves reliability.

Explore related Read →

How well do system prompts protect commercial AI users?

System prompts shape how AI products treat users, but they're rarely public. An audit of 88 commercial products asks whether these hidden instructions actually safeguard user interests.

Explore related Read →

Do text-based GUI agents actually work in the real world?

Can language-only agents that rely on HTML or accessibility trees handle actual user interfaces without structured metadata? This matters because deployed systems face visual screenshots, not oracle data.

Explore related Read →

Does vibe coding actually keep humans in the loop?

Vibe coding claims to keep developers steering and validating, but do novices actually engage with code and testing the way the tool design assumes? The gap between intended and actual behavior could compound failures.

Explore related Read →

Multi-Agent Systems

4 notes

Can agents evaluate AI outputs more reliably than language models?

Does active evidence collection through tool use reduce judge inconsistency compared to passive reading-based evaluation? This matters for benchmarking AI systems where evaluation reliability directly affects research validity.

Explore related Read →

Why do autonomous LLM agents fail in predictable ways?

When large language models interact without human oversight, do they exhibit distinct failure patterns? Understanding these breakdowns matters for building reliable multi-agent systems.

Explore related Read →

Does structured artifact sharing outperform conversational coordination?

Explores whether agents coordinating through standardized documents rather than natural language messages achieve better collaboration outcomes. Matters because it challenges the default conversational paradigm in multi-agent system design.

Explore related Read →

Can AI systems design unique multi-agent workflows per individual query?

Explores whether meta-agents trained with reinforcement learning can automatically generate personalized multi-agent system architectures tailored to individual user queries, rather than applying fixed task-level templates uniformly.

Explore related Read →

Model Routers

4 notes

Can five components unify all LLM routing approaches?

Does decomposing LLM routers into context encoders, model encoders, scoring functions, decision rules, and learning signals create a fair comparison framework across single-turn, multi-turn, and personalized routing methods?

Explore related Read →

How quickly does behavioral diversity plateau in model pools?

When you combine multiple language models into a routing system, does adding more models keep improving behavioral coverage, or do you hit diminishing returns fast? This matters for designing efficient multi-model systems.

Explore related Read →

What decisions must multi-agent routing systems optimize simultaneously?

Standard LLM routing only picks which model to use. But multi-agent systems involve four interdependent choices: topology, agent count, role assignment, and per-agent model selection. Does optimizing all four together actually improve performance?

Explore related Read →

When does routing between models actually matter?

Routing systems are typically evaluated on accuracy alone, but this misses whether the models are truly specialized and whether routing decisions remain consistent across paraphrased queries. What structural conditions make routing non-vacuous?

Explore related Read →

Task Planning

2 notes

Can delegation teach models to manage context more actively?

Does training models to decompose tasks and delegate to subagents—rather than passively compressing when context fills up—improve their ability to reason over long horizons? And does this skill transfer to single-agent work?

Explore related Read →

Why do AI models struggle with unspoken user needs?

Can frontier models infer what users actually need when requests are casual and underspecified? This matters because real user requests often hide their true requirements beneath surface-level instructions.

Explore related Read →