SYNTHESIS NOTE
Topics›Reinforcement Learning›this note

Does reinforcement learning update only a small fraction of parameters?

Investigating whether RL algorithms consistently modify only 5–30% of model parameters across different LLMs and RL methods, and what structural properties those sparse updates possess.

Synthesis note · 2026-02-22 · sourced from Reinforcement Learning

The surprising finding is not that RL changes models — it's how little it changes them. Across PPO, GRPO, DPO, and four other algorithms applied to ten different LLM families, RL consistently updates only 5-30% of parameters. The rest remain effectively unchanged. This sparsity is intrinsic — no explicit sparsity-promoting regularizations or architectural constraints are applied.

The critical nuance is that these sparse updates are nearly full-rank. This is not low-rank adaptation (as in LoRA). The updated parameters span almost the full subspace that the parameter matrices can represent. So RL selects a small subset of parameters, but that subset is geometrically rich enough to represent complex transformations. The distinction matters: low-rank would mean RL operates in a constrained subspace; sparse-but-full-rank means RL identifies which parameters matter while preserving full expressivity.

Three additional properties make this pattern robust. First, subnetworks identified from different random seeds show substantially greater overlap than chance, suggesting the subnetwork is a structural property of the pretrained model, not an artifact of training. Second, finetuning the subnetwork alone recovers both the test accuracy and the actual parameter values of full finetuning. Third, the sparsity is distributed — nearly all parameter matrices receive similarly sparse updates rather than concentrating in a subset of layers.

The authors conjecture this sparsity arises primarily from training on data near the policy distribution. Since Does RL improve domain reasoning by adding knowledge or removing it?, the sparse-but-full-rank pattern provides a mechanistic explanation: RL doesn't need to transform the entire model because most of the model is already adequate. It just needs to adjust a targeted subset — the parameters that control which reasoning paths are taken.

This has implications for efficient RL training. If the effective parameter footprint is 5-30%, techniques that exploit this sparsity (targeted updates, efficient memory use) could dramatically reduce RL training cost without sacrificing quality.

Token-level 80/20 parallel: The parameter-level sparsity has a striking token-level analog. The "Beyond 80/20" analysis of RLVR shows that high-entropy minority tokens — the ~20% of tokens where the model is most uncertain — are the critical forking points that carry most of the learning signal. Restricting gradient updates to only these 20% of tokens matches or exceeds full-token updates (+11.04 on AIME'25 for Qwen3-32B). The remaining 80% of tokens are low-entropy, already-decided outputs where gradient updates add noise rather than signal. This creates a dual sparsity picture: RL updates 5-30% of parameters, and the effective signal comes from ~20% of tokens. Both forms of sparsity are intrinsic — not imposed by regularization — and both suggest RL is fundamentally a targeted refinement process rather than a wholesale model transformation. See Do high-entropy tokens drive reasoning model improvements?.

The same sparse-update structure appears in SFT. Core Parameter Isolation Fine-Tuning (CPI-FT) identifies task-specific "core parameter regions" — the parameters with largest update magnitudes during individual task fine-tuning — and shows that these regions are concentrated and task-specific. CPI-FT exploits this by transplanting core parameters from individually fine-tuned models and SLERP-merging non-core parameters, consistently outperforming full multi-task SFT. The key finding: full multi-task SFT (uniform parameter updates across all tasks) is consistently the worst performer — temporal task scheduling alone is insufficient without explicit structural parameter isolation. This extends the RL sparsity finding to supervised fine-tuning: task-relevant changes naturally concentrate in specific parameter regions regardless of whether the training signal is reward-based or loss-based. See Can isolating task-specific parameters prevent multi-task fine-tuning interference?.

Inquiring lines that read this note 117

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Why do stronger reasoning capabilities create tradeoffs with instruction following? How do pretraining biases affect reward signal effectiveness in RLVR? Does RL create genuinely new reasoning capabilities or refine existing ones? How do capability benchmark scores systematically misrepresent true model abilities? Can self-generated feedback reliably guide model training without ground truth? What compositional reasoning failures limit large language models despite scale? How do surface patterns enable correct outputs but reduce robustness? Do language models develop actual world models or merely task heuristics? How does AI adoption across firms reshape employment and inequality? What training dynamics and scale trigger emergence of reasoning capabilities? How should inference compute be allocated based on problem difficulty? How does policy entropy collapse constrain scaling of reasoning-focused RL? How should designers communicate what AI systems truly are and can do? How much do training data properties shape model reasoning? What capability trade-offs arise from domain specialization through fine-tuning? How should agents manage memory granularity to improve long-term performance? What makes step-level supervision effective for complex reasoning traces? Can diffusion models match autoregressive performance on language generation tasks? Can preference-based training achieve better behavior optimization than supervised fine-tuning alone? What role does sparsity play in model behavior and scaling decisions? How can conversational agents maintain consistent personas across multi-turn dialogue? Can brute-force automated research substitute for iterative depth and human research intuition? Why does adding new knowledge through fine-tuning degrade existing capabilities? What trajectory-level metrics beyond task success best evaluate agent performance? Can inoculation prompting prevent emergent misalignment after reward hacking? Can inference-time compute effectively substitute for model scale? How does harness optimization generalize across different model architectures and domains? How does improved reasoning affect models' ability to acknowledge uncertainty? Can iterative DPO replicate online reinforcement learning dynamics for research? What fundamental constraints limit how effectively agents can improve themselves? Can we reliably detect when models game evaluations? Can compression size predict model complexity better than parameter count alone? How do agent-learned skills transfer and improve across different tasks?

Related concepts in this collection 7

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
19 direct connections · 162 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

rl updates only 5-30 percent of parameters in sparse but full-rank subnetworks