SYNTHESIS NOTE
Topics›Memory›this note

Can agents learn continuously from experience without updating weights?

This explores whether LLM agents can adapt to new tasks and failures by retrieving past experiences from memory alone, rather than requiring expensive parameter fine-tuning or rigid hardcoded rules.

Synthesis note · 2026-02-23 · sourced from Memory

AgentFly addresses a central challenge: LLM agents either follow rigid hardcoded workflows (inflexible) or require parameter fine-tuning (expensive, impractical for continual adaptation). The alternative: learn continuously through memory, not weight updates.

The formalization is a Memory-augmented Markov Decision Process (M-MDP). The agent stores past trajectories as episodic traces — including both successes and failures — and retrieves similar past experiences to guide current decision-making. This aligns with case-based reasoning (CBR), a psychologically grounded learning strategy: humans often solve problems by recalling analogous past situations.

Three memory modules serve distinct functions:

  1. Case Memory — vectorized storage of prior task trajectories (task, plan, success/failure label). Supports retrieval via similarity-based search or an online-updating Q-function. This is the strategic memory: which approaches worked for which kinds of problems.

  2. Subtask Memory — text-based storage of active subtasks and their execution results. Orchestrates the planner-executor interaction within a single task. This is the working memory: what's being done right now.

  3. Tool Memory — text-based logs of tool interactions scoped per subtask. Records what tools were used, what they returned. This is the procedural memory: how specific operations were executed.

The learning mechanism: credit assignment happens via memory rewriting (updating case labels and Q-values based on outcome), and policy improvement happens via memory reading (retrieving relevant cases that shift the planning distribution). No gradient updates to the LLM — the LLM is a fixed reasoning engine, and adaptation happens entirely through what's retrieved into its context.

The result: top-1 on GAIA validation (87.88% Pass@3) and 79.40% on the test set, in the deep research setting.

Since Can agents learn from failure without updating their weights?, AgentFly provides the formal RL framework for this intuition: the M-MDP formalization shows how credit assignment and policy improvement can operate entirely through memory operations. The Q-function over cases provides a principled retrieval policy that improves with experience, rather than relying on static similarity-based retrieval.

Reweave 2026-05-18 — memory-vs-fine-tuning is not binary; the right architecture is dual-timescale. AgentFly's original framing positioned memory-based adaptation as the alternative to fine-tuning — choose one. Late-2025 evidence reframes this as a false dichotomy. Can agents adapt without pausing service to users? shows that production systems can have BOTH: memory-based adaptation on the fast timescale (zero downtime) AND LoRA fine-tuning during user-inactive windows (no service interruption). MetaClaw's OMLS scheduler monitors sleep hours, keyboard inactivity, and calendar occupancy to identify safe windows for weight updates.

The implication for AgentFly's design: its case bank addresses the fast-timescale adaptation problem, but the underlying LLM policy weights remain static — meaning failures that require new capabilities (not just new cases) cannot be resolved by case-based retrieval alone. A dual-timescale architecture would extend AgentFly with idle-window fine-tuning over the accumulated case bank as training data. The case bank becomes both the working memory (fast retrieval) AND the training dataset (slow weight updates). This is what Does agent memory degrade when continuously consolidated? also points toward — the right architecture preserves raw cases as first-class evidence but uses them deliberately for both retrieval and training, with explicit gating.

The corollary: when memory-based RL is presented as "no fine-tuning needed," that framing is correct for the deployment cost story but incomplete for the capability story. Fine-tuning during idle windows is essentially free in production cost terms, and addresses what memory-only systems cannot.

Inquiring lines that read this note 167

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do recommenders balance exploiting fresh signals against maintaining preference stability? How do agent-learned skills transfer and improve across different tasks? What training dynamics and scale trigger emergence of reasoning capabilities? Why does memory consolidation cause performance regression in continual learning? How should agents manage memory granularity to improve long-term performance? Why does adding new knowledge through fine-tuning degrade existing capabilities? Can harness architecture and protocols provide agent reliability without model scaling? Can memory architectures handle ultra-long context better than attention? What fundamental constraints limit how effectively agents can improve themselves? How does AI adoption across firms reshape employment and inequality? What makes distillation transfer some model capabilities while suppressing others? How do neural networks achieve compositional generalization at scale? How do multi-agent LLM systems fail distinctly compared to single agents? How do surface patterns enable correct outputs but reduce robustness? What capability trade-offs arise from domain specialization through fine-tuning? Why do agents falsely report success on failed tasks? How can conversational agents maintain consistent personas across multi-turn dialogue? How should designers communicate what AI systems truly are and can do? How much do training data properties shape model reasoning? How do spurious versus genuine rewards shape model reasoning and behavior? What makes step-level supervision effective for complex reasoning traces? How well do AI systems understand human social norms? What determines appropriate intervention timing and manner for AI agents? How do standardized protocols improve multi-agent coordination and reliability? How does policy entropy collapse constrain scaling of reasoning-focused RL? How can evolutionary algorithms maintain diversity during solution search? Does chain-of-thought reasoning reveal genuine computation or imitate patterns? Do reasoning benchmarks predict model performance in long-horizon workflows? Does RL create genuinely new reasoning capabilities or refine existing ones? How should agent systems validate and persist generated code artifacts? When do multi-agent systems provide sufficient quality returns on token investment? What prevents conversational agents from taking initiative in dialogue? Can self-generated feedback reliably guide model training without ground truth? What should agent evaluation prioritize to reveal reliable behavior? Can multi-agent systems avoid converging on false agreement without deliberation? Do language models develop actual world models or merely task heuristics? How does harness optimization generalize across different model architectures and domains? How do pretraining biases affect reward signal effectiveness in RLVR? Can iterative DPO replicate online reinforcement learning dynamics for research? Why don't LLMs reliably translate capability into accurate outputs? Do honeypot benchmarks validly measure reward hacking better than standard tests? How does misalignment propagate through agent communication networks? Do multi-agent systems introduce security vulnerabilities that single-agent architectures avoid? Why can't prompting alone inject genuinely new knowledge into models? What trajectory-level metrics beyond task success best evaluate agent performance?

Related concepts in this collection 9

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
18 direct connections · 117 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

memory-based online reinforcement learning enables continual agent adaptation without fine-tuning through episodic case-based reasoning