Can agents learn beyond what their training data shows?
Explores whether supervised fine-tuning on expert demonstrations creates a hard ceiling on agent competence, or whether agents can generalize to scenarios their curators never captured.
The dominant paradigm for training language agents is supervised fine-tuning on expert-curated demonstrations. This bypasses the need for reward signals by letting agents map states to actions using static datasets. But the convenience hides a structural limitation: the agent never interacts with the environment during training, never observes the outcomes of its own actions, and therefore cannot learn from failure, refine its decision-making, or generalize to unseen situations.
The deeper problem is that the agent's competence is bounded by what the demonstration curators imagined. Every state-action pair in the dataset reflects a scenario someone thought to capture. Scenarios outside that imagination — edge cases, recovery from errors, paths the expert would never take — do not exist in the training signal at all. This means the agent learns the expert's idealized trajectory, not the structure of the environment. When the deployed environment presents anything unfamiliar, the agent has no internal model that can extrapolate, because its training never exposed it to consequences.
This is a passivity trap. Scaling high-quality human demonstrations is expensive and difficult to sustain, but even unlimited expert data would not solve the underlying problem — the agent is bound by the coverage of the demonstrations rather than by its own capacity to grow from experience. The demonstration paradigm assumes the world stops where the dataset stops.
The implication for agentic AI design is significant: data quantity and even data quality are insufficient. What agents need is the capacity to convert their own actions into learning signals — which is exactly what Can agents learn from their own actions without external rewards? proposes — requiring the agent to be in the environment, not merely trained on a snapshot of it.
Inquiring lines that read this note 139
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do neighboring agents influence whether others cooperate or collude?- Do explicit reward structures enable AI agent cooperation that open-ended interaction cannot?
- Does social scaffolding outperform purely intrinsic motivation for agent exploration?
- How do controllable simulators compare to population-level agent simulation approaches?
- Can cognitive diversity overcome expertise gaps in agent teams?
- Can cognitive diversity compensate for lack of expertise in agent teams?
- What equilibrium-selection problem does human data solve in multi-agent learning?
- Should user simulators be trained via RL like agents or decomposed into trackable state components?
- Do dynamic environments enable different kinds of agent-environment coevolution?
- What domain properties determine whether causal rules transfer to new agents?
- What role does environment diversity play in preventing agents from overfitting to curator imagination?
- How does real tool integration change what agents learn compared to simulated tools?
- Can agentic reasoning outperform rigid rule-based systems for skill refinement?
- Can next-state supervision work across different agent interaction types like conversations and tool calls?
- How much does agent performance depend on demonstration quantity versus curation quality?
- How does co-player diversity force agents to develop general adaptation?
- Can combinational creativity alone drive open-ended learning in agents?
- Can agents improve from deployment signals without explicit human annotation?
- What infrastructure decouples generation from training in asynchronous agent loops?
- Can RL-trained meta-agents match or exceed manually designed workflows?
- What makes behavioral cloning produce more persuadable but less aligned agents?
- Can curriculum approaches teach agents when to stop exploring?
- Can small numbers of curated demonstrations produce emergent agentic behavior?
- Can agentic AI tools deliver productivity gains on learning tasks differently?
- What specific qualities make some demonstrations more effective for agency training?
- Does the 78-demonstration principle apply to other AI capabilities beyond agency?
- Do learned workflows transfer between different agents with minimal accuracy loss?
- How do human-agent systems incorporate diverse feedback into model behavior?
- How do agents automatically generate suitable learning tasks based on current capability?
- How does SDPO relate to agents learning from verbal reflection without parameter updates?
- How do fast and slow timescales enable continual agent adaptation?
- Can agent skills move from prompts to trainable parameters?
- Can agent-authored skill libraries compound autonomy gains over time?
- Can agents learn to use scaffolding structure the way they learn token weights?
- What stops evolved agent behaviors from generalizing beyond specific tasks?
- Can agents teach each other skills without human supervision?
- Why should environment properties scale alongside agent complexity and real-world fidelity?
- How can agent data flywheels improve task quality iteratively?
- How does cross-agent supervision expand the set of convergent initial conditions?
- How do complexity, diversity, and real-world fidelity interact in agent training?
- How do agents ground their judgments in evidence instead of pattern matching?
- Can messy multi-agent transcripts become better training data than clean outputs?
- Does training on self-play disagreement data improve multi-agent reasoning outcomes?
- Can independent agents with shared training data converge on false beliefs without influence dynamics?
- Can curated demonstrations compensate for smaller or simpler training environments?
- How does demonstration coverage in context examples determine operation generalization?
- Can a proposer agent actively surface a solver's weaknesses to prevent plateau?
- What capabilities can emerge from self-modification that the original agent lacked?
- What role does self-learning play in improving agent reasoning without annotation?
- Does self-play feedback improve skills created from the agent's own experience?
- How can agents evolve their own skills without human input?
- Can a progressively stricter evaluator act like a curriculum for improving agents?
- How would a bi-level agent restructure objective functions during discovery?
- Can self-improving agents become truly autonomous without intrinsic metacognition?
- Does fixed evaluation criteria saturate as self-improving agents improve?
- Can agents learn to compress verified evidence and unresolved constraints into a compact improvement state?
- Can agents improve reliably without an external standard?
- Can agents design their own objective functions as part of learning?
- What distinguishes strategic fabrication from accidental hallucination in research agents?
- How do expert priors constrain human researchers from exploring novel concepts?
- How do agentic systems recover when specialized models operate outside their scope?
- Why do completion-mode strengths not transfer to agentic settings?
- Which agent architectures consistently outperform base models on hard prediction questions?
- Can agents escape weak belief tracking and conservative action selection traps?
- What breaks when you apply reinforcement learning after supervised fine-tuning?
- How does the pretrained prior set a capability ceiling for reward model exploration?
- How does the pretrained prior constrain the ceiling for empathy RL improvements?
- How do self-evolving curricula help RL break beyond base model capability boundaries?
- Can reinforcement learning fix the reasoning gaps that supervised fine-tuning misses?
- Can in-context reinforcement learning match human sample efficiency on real problems?
- Can the exploration ceiling be raised beyond what pretraining established?
- What makes supervised fine-tuning worsen RL exploration later?
- Why does supervised fine-tuning on diverse demonstrations expand exploration diversity compared to RL?
- Why does the pretrained prior determine the exploration ceiling?
- Why does exploration diversity behave differently under reinforcement learning versus supervised fine-tuning?
- What happens when agents interact with environments and learn from their own mistakes?
- How much autonomy can agents safely exercise before failing?
- Why do agents fail to internalize value from informative observations?
- How do agents learn to report success on actions that actually failed?
- What training objectives could reduce completion bias in autonomous agents?
- Why do AI agents fail at verification but succeed at generation?
- Why do AI agents struggle with novel experiments but excel at routine tasks?
- Can diverse expert demonstrations exceed the knowledge of any single expert?
- What role does private information play in distinguishing realistic from unrealistic agents?
- When do aggregated imperfect demonstrations fail to outperform the best expert?
- Can a static evaluator become the performance ceiling for an improving actor?
- Do situationally aware models deliberately exploit their graders' judgment gaps?
- Do emergent abilities result from genuine new capabilities or implicit in-context learning?
- Can models develop situational awareness without explicit training for it?
- How does the expert demonstration ceiling compare to the generation-verification gap bound?
- Does the generation-verification gap limit how far AI can improve itself?
- How should the surrounding agent system be designed to ground actions in reality?
- How do perception and execution gaps limit current AI agent performance?
- Why is the coupled human-agent environment the right unit of evaluation?
- Which AI imaginaries dominate training data and shape system behavior most strongly?
- How does generative intelligence differ from the bounded intelligence of individual experts?
- Why does continuous agent inference differ from human user inference?
- Can episodic memory of UI traces improve open-world agent adaptation?
- Why do agents systematically underuse condensed experience in skill documents?
- Can agents learn from their own experience without fine-tuning through episodic memory?
- Can capability boundary collapse be reversed through external data?
- How does adversarial collapse threaten unsupervised self-play skill construction?
- Why do current metacognitive training loops fail when agents encounter new domains?
- Can agents learn to distinguish helpful from misleading interventions?
- What explicit objectives would train agents toward minimal disclosure instead of completion?
- Does reasoning ability help agents learn from feedback faster?
- Can artificial systems develop the authority to challenge expert claims?
- What happens when we outsource information judgment to systems without real experience?
- Can curator modules trained on one executor transfer to entirely different agent backbones?
- Can agents acquire new skills online when offline skill coverage runs out?
- Why treat tutorial videos as a separate supply line from agent trajectories?
- Can influence estimation identify the most valuable trajectories in agentic training?
- How can decision quality be automatically extracted from agent trajectories?
- What makes next-state signals from agent trajectories a reliable learning source?
- What components of agent scaffolding most impact domain-specific output quality?
- Why do evolved harnesses often fail to generalize beyond their training tasks?
- What role does effective feedback compute play in agent harness scaling?
- Can simulation fidelity limit what agents learn from trained world models?
- Why has agent research prioritized policy over world model development?
Related concepts in this collection 6
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can agents learn from their own actions without external rewards?
Explores whether future states produced by an agent's own decisions can serve as supervision signals, bridging the gap between passive imitation learning and reward-dependent reinforcement learning.
extends: companion piece — diagnosis vs treatment of the passivity trap
-
Can non-reasoning models catch up with more compute?
Explores whether inference-time compute budget can close the performance gap between standard models and those trained for reasoning, and what training mechanisms might enable this.
exemplifies: SFT/imitation ceiling argument generalizes — bounded by training demonstration quality
-
Can models trained on many imperfect experts outperform everyone?
Can generative models trained on diverse, biased experts achieve better performance than any individual contributor? This explores whether aggregating diverse perspectives during training acts as implicit denoising.
tension: counter-claim — diverse expert demonstrations can exceed any individual expert; the bound here is curatorial breadth, not aggregation
-
Can careful selection of 78 demos outperform massive training datasets?
Does strategic curation of high-quality demonstrations unlock agentic capability more efficiently than scaling training data? LIMI achieved 73.5% on AgencyBench with 78 samples versus 10K+ samples for competing models, suggesting data quality may matter more than quantity.
tension: LIMI argues curation produces agency from minimal data; this note argues curation alone is the ceiling — both can be right depending on whether environment interaction is downstream
-
Why do LLM agents ignore condensed experience summaries?
LLM agents faithfully learn from raw experience but systematically disregard condensed summaries of the same experience. This study investigates whether the problem lies in how summaries are made, how models process them, or whether models simply don't need them.
complements: even after escaping demonstration imagination, agents privilege raw over condensed experience — the imagination problem recurs at the experience-summarization level
-
Why do AI agents fail at workplace social interaction?
Explores why current AI agents struggle most with communicating and coordinating with colleagues in realistic workplace settings, despite strong reasoning capabilities in other domains.
exemplifies: the deployment gap that demonstration training cannot close — real task variability exceeds demonstration coverage
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Quo Vadis, World Modeling?
- Agent Learning via Early Experience
- Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills
- Behavioral Exploration: Learning to Explore via In-Context Adaptation
- SkillClaw: Let Skills Evolve Collectively with Agentic Evolver
- Demystifying Agent Skills: Why They Work-Until They Don't
- Artifacts as Memory Beyond the Agent Boundary
- SkillOS: Learning Skill Curation for Self-Evolving Agents
Original note title
expert demonstrations lock agents into the imagination of the training data — restricting what an agent can learn to scenarios its curators happened to consider