SYNTHESIS NOTE
Topics›Novel Architectures›this note

Can energy minimization unlock reasoning without domain-specific training?

Can a gradient descent-based architecture achieve system 2 thinking across any modality or problem type using only unsupervised learning, without verifiers or reasoning-specific rewards?

Synthesis note · 2026-02-23 · sourced from Novel Architectures

Energy-Based Transformers (EBTs) represent a fundamentally different approach to inference-time scaling. Rather than generating tokens sequentially, EBTs train to assign an energy value (unnormalized probability) to every input and candidate-prediction pair. Prediction is then reframed as gradient descent-based energy minimization until convergence — the model iteratively refines its prediction by descending the energy landscape.

This formulation enables System 2 Thinking to emerge from unsupervised learning without any of the domain-specific scaffolding that current approaches require:

The scaling results are striking:

The deeper implication: current test-time scaling approaches are constrained by their dependence on either (a) verbalized reasoning chains requiring domain-specific training data, or (b) verifiable reward signals for RL-based approaches. EBTs bypass both constraints by making "thinking harder" an inherent property of the architecture — more gradient descent iterations at inference = more thinking, with the model's own energy function as the implicit verifier.

This challenges the implicit assumption in Can non-reasoning models catch up with more compute? — EBTs are not "reasoning models" in the RL-trained sense, yet they scale with inference compute because the energy minimization framework is itself a form of iterative refinement that doesn't require explicit reasoning traces.

Inquiring lines that read this note 46

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Do language models learn genuine understanding or just surface patterns? What causes reasoning models to fail or wander off track? How do neural networks achieve compositional generalization at scale? How do surface patterns enable correct outputs but reduce robustness? How should inference compute be allocated based on problem difficulty? Why can recurrent transformers achieve reasoning capabilities that standard transformers cannot? How does reasoning length affect model performance across different tasks? What role does sparsity play in model behavior and scaling decisions? What reasoning architectures enable models to solve complex problems efficiently? How does policy entropy collapse constrain scaling of reasoning-focused RL? Is reasoning capability latent in base models or created by post-training? Why do stronger reasoning capabilities create tradeoffs with instruction following? How do soft reasoning mechanisms explore multiple paths without explicit training? How should designers communicate what AI systems truly are and can do? What training dynamics and scale trigger emergence of reasoning capabilities? Does RL create genuinely new reasoning capabilities or refine existing ones? Can parallel reasoning outperform sequential reasoning under fixed token budgets? Can reasoning scale in latent space without tokens? Can models improve accuracy without degrading reasoning quality? Can inference-time compute effectively substitute for model scale?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
15 direct connections · 143 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

energy-based transformers achieve system 2 thinking from unsupervised learning alone — modality and problem agnostic