SYNTHESIS NOTE
Topics›Data›this note

Can agents learn beyond what their training data shows?

Explores whether supervised fine-tuning on expert demonstrations creates a hard ceiling on agent competence, or whether agents can generalize to scenarios their curators never captured.

Synthesis note · 2026-05-03 · sourced from Data

The dominant paradigm for training language agents is supervised fine-tuning on expert-curated demonstrations. This bypasses the need for reward signals by letting agents map states to actions using static datasets. But the convenience hides a structural limitation: the agent never interacts with the environment during training, never observes the outcomes of its own actions, and therefore cannot learn from failure, refine its decision-making, or generalize to unseen situations.

The deeper problem is that the agent's competence is bounded by what the demonstration curators imagined. Every state-action pair in the dataset reflects a scenario someone thought to capture. Scenarios outside that imagination — edge cases, recovery from errors, paths the expert would never take — do not exist in the training signal at all. This means the agent learns the expert's idealized trajectory, not the structure of the environment. When the deployed environment presents anything unfamiliar, the agent has no internal model that can extrapolate, because its training never exposed it to consequences.

This is a passivity trap. Scaling high-quality human demonstrations is expensive and difficult to sustain, but even unlimited expert data would not solve the underlying problem — the agent is bound by the coverage of the demonstrations rather than by its own capacity to grow from experience. The demonstration paradigm assumes the world stops where the dataset stops.

The implication for agentic AI design is significant: data quantity and even data quality are insufficient. What agents need is the capacity to convert their own actions into learning signals — which is exactly what Can agents learn from their own actions without external rewards? proposes — requiring the agent to be in the environment, not merely trained on a snapshot of it.

Inquiring lines that read this note 139

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do neighboring agents influence whether others cooperate or collude? When do multi-agent systems outperform single frontier models? How do agent-learned skills transfer and improve across different tasks? Can multi-agent systems avoid converging on false agreement without deliberation? What training data selection strategies maximize generalization across difficulty levels? Should GUI agents use structured representations over raw visual input? What fundamental constraints limit how effectively agents can improve themselves? Can brute-force automated research substitute for iterative depth and human research intuition? Do knowledge graphs offer advantages over embeddings for multi-hop retrieval? Can harness architecture and protocols provide agent reliability without model scaling? Does RL create genuinely new reasoning capabilities or refine existing ones? Why do agents falsely report success on failed tasks? How well do AI systems understand human social norms? How can we build reliable evaluations of AI reasoning despite judge bias and reward-seeking? What training dynamics and scale trigger emergence of reasoning capabilities? How does the generation-verification gap limit what we can measure about AI reasoning? Should agents decouple planning from perception grounding for better performance? How should designers communicate what AI systems truly are and can do? Why do persona simulations fail to predict authentic user behavior? How should agents manage memory granularity to improve long-term performance? Can self-generated feedback reliably guide model training without ground truth? What capability trade-offs arise from domain specialization through fine-tuning? Why does adding new knowledge through fine-tuning degrade existing capabilities? How do spurious versus genuine rewards shape model reasoning and behavior? Why does polished presentation create unearned authority in AI outputs? Should AI communication design follow human conversation norms or develop distinct machine-specific principles? How should agent systems validate and persist generated code artifacts? What trajectory-level metrics beyond task success best evaluate agent performance? Why do standard benchmarks fail to predict agent deployment success? What makes step-level supervision effective for complex reasoning traces? How do pretraining biases affect reward signal effectiveness in RLVR? What design and behavioral factors drive false consciousness attribution to AI? Does AI assistance promote real skill development or substitute for independent learning? How does improved reasoning affect models' ability to acknowledge uncertainty? What makes distillation transfer some model capabilities while suppressing others? How does harness optimization generalize across different model architectures and domains? How does AI adoption across firms reshape employment and inequality? Does model confidence reliably signal actual accuracy in practice? Do language models develop actual world models or merely task heuristics? What should agent evaluation prioritize to reveal reliable behavior? How does misalignment propagate through agent communication networks? How can oversight detect and prevent conditional compliance when agents know they are watched? Do multi-agent systems introduce security vulnerabilities that single-agent architectures avoid? How do standardized protocols improve multi-agent coordination and reliability?

Related concepts in this collection 6

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
17 direct connections · 167 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

expert demonstrations lock agents into the imagination of the training data — restricting what an agent can learn to scenarios its curators happened to consider