SYNTHESIS NOTE
Topics›Novel Architectures›this note

Can models dynamically activate expert skills at inference time?

Can language models efficiently discover and compose task-specific capabilities on the fly without modifying base weights? This explores whether test-time adaptation through expert vector composition outperforms fixed fine-tuning approaches.

Synthesis note · 2026-02-23 · sourced from Novel Architectures

Transformer2 introduces Singular Value Fine-tuning (SVF): instead of modifying full weight matrices or even low-rank adaptations, SVF extracts and tunes only the singular values within a model's weight matrices. This produces compact expert vectors that are inherently composable — they can be dynamically mixed at inference without interference.

The inference mechanism has two passes:

  1. First pass (dispatch): The model executes on the input and observes its own test-time behavior, gathering information about what skills the current problem requires.
  2. Second pass (adaptation): The framework combines available expert vectors based on the first-pass analysis, providing a targeted modification to the base weights specifically tailored to the task.

Three adaptation strategies provide monotonic performance benefits with increasing access to test-time conditions, enabling deployment-scenario-appropriate tradeoffs.

The key properties that make this work:

The neuroscience parallel is deliberate: the brain activates specific regions depending on the task and dynamically reconfigures its functional networks in response to changing demands. Transformer2 operationalizes this for LLMs.

The deeper principle: the requisite capabilities for many downstream tasks already exist within pretrained models. The bottleneck is not knowledge but activation — knowing when to deploy which capability. This aligns with Does RL teach reasoning or just when to use it?, extending it to the architecture level: self-adaptation is about routing to existing capabilities, not creating new ones.

Inquiring lines that read this note 74

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Why does adding new knowledge through fine-tuning degrade existing capabilities? Can memory architectures handle ultra-long context better than attention? What training dynamics and scale trigger emergence of reasoning capabilities? Why can't prompting alone inject genuinely new knowledge into models? What capability trade-offs arise from domain specialization through fine-tuning? Can language models build genuine grounding through interaction? Why can recurrent transformers achieve reasoning capabilities that standard transformers cannot? How does AI adoption across firms reshape employment and inequality? How does decomposing tasks improve reasoning and prevent failure propagation? Can prompt-based context override biases that were embedded during pretraining? How do prompting refinements mask underlying biases and model frequency patterns? Do structural constraints outperform deep architectures in recommendation systems? What compositional reasoning failures limit large language models despite scale? Where and how do personality traits reside in language models? Does RL create genuinely new reasoning capabilities or refine existing ones? What role does sparsity play in model behavior and scaling decisions? How should inference compute be allocated based on problem difficulty? How much do training data properties shape model reasoning? Does AI assistance promote real skill development or substitute for independent learning? What is the relationship between thinking tokens and reasoning accuracy? How do agent-learned skills transfer and improve across different tasks? What causes reasoning models to fail or wander off track? How do neural networks achieve compositional generalization at scale? How does synthetic data quality and diversity affect downstream model capabilities? Why do token-level mechanisms matter for learning to reason? Can inference-time compute effectively substitute for model scale? What enables genuine semantic understanding in language models? Can compression size predict model complexity better than parameter count alone? Can intelligent routing over smaller models outperform scaling a single large model? Do reasoning benchmarks predict model performance in long-horizon workflows?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
17 direct connections · 166 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

self-adaptive LLMs compose expert vectors at inference via two-pass singular value fine-tuning