Can energy minimization unlock reasoning without domain-specific training?
Can a gradient descent-based architecture achieve system 2 thinking across any modality or problem type using only unsupervised learning, without verifiers or reasoning-specific rewards?
Energy-Based Transformers (EBTs) represent a fundamentally different approach to inference-time scaling. Rather than generating tokens sequentially, EBTs train to assign an energy value (unnormalized probability) to every input and candidate-prediction pair. Prediction is then reframed as gradient descent-based energy minimization until convergence — the model iteratively refines its prediction by descending the energy landscape.
This formulation enables System 2 Thinking to emerge from unsupervised learning without any of the domain-specific scaffolding that current approaches require:
- No modality restrictions (works on both text and images)
- No problem-specific design (not limited to verifiable domains like math/code)
- No additional supervision beyond unsupervised pretraining (no verifiers, no verifiable rewards)
The scaling results are striking:
- Training: Up to 35% higher scaling rate than Transformer++ with respect to data, batch size, parameters, FLOPs, and depth
- Inference: 29% more improvement from additional test-time compute on language tasks than Transformer++
- Generalization: Larger performance improvements on data farther out-of-distribution — suggesting EBTs generalize better than existing approaches
- Efficiency: Outperform Diffusion Transformers on image denoising with fewer forward passes
The deeper implication: current test-time scaling approaches are constrained by their dependence on either (a) verbalized reasoning chains requiring domain-specific training data, or (b) verifiable reward signals for RL-based approaches. EBTs bypass both constraints by making "thinking harder" an inherent property of the architecture — more gradient descent iterations at inference = more thinking, with the model's own energy function as the implicit verifier.
This challenges the implicit assumption in Can non-reasoning models catch up with more compute? — EBTs are not "reasoning models" in the RL-trained sense, yet they scale with inference compute because the energy minimization framework is itself a form of iterative refinement that doesn't require explicit reasoning traces.
Inquiring lines that read this note 46
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Do language models learn genuine understanding or just surface patterns? What causes reasoning models to fail or wander off track?- Can surface heuristics override implicit constraints in domain-specific reasoning?
- Can explicit optimal algorithms prevent reasoning model collapse at high complexity?
- Why do human-designed neural architectures eventually get replaced by learned ones?
- What inductive bias would force models to learn Newtonian mechanics instead of shortcuts?
- How do classical mechanics and statistical mechanics provide methodological templates for learning theory?
- How can neural networks be interpretable by design rather than post-hoc?
- How much does domain shift limit the mechanisms a bilevel system can autonomously discover?
- Can gradient approximation at equilibrium replace backpropagation through time in practice?
- Can latent recurrence and energy minimization both escape the same computational depth constraints?
- Can deterministic recurrent depth achieve the computational benefits of stochastic reasoning?
- Can energy-based transformers achieve deep reasoning without supervision?
- Can looped architectures achieve reasoning abilities that fixed-depth models cannot?
- How do gradient descent iterations at inference compare to chain-of-thought reasoning chains?
- Why does extended chain-of-thought reasoning fail to improve numerical optimization performance?
- Can we transfer reasoning structure without copying surface form?
- How do beam search and MCTS traverse reasoning topologies?
- Can a single architecture represent both physical and mental possibility spaces?
- Can targeted activation steering surface latent reasoning in base models?
- How does factoring perception from reasoning improve sparse-label learning?
- Why do recursive belief models require different training than logical derivation?
- What makes thought identifiability provable without auxiliary training data?
- Do base models contain latent reasoning that minimal training can unlock?
- Can models reason at inference without specialized internal training?
- How much training data is truly necessary to unlock latent model reasoning?
- Can distillation from stronger models create genuinely new reasoning abilities?
- Can minimal training signals unlock latent reasoning capability in base models?
- Can minimal training signals unlock reasoning already latent in pretrained representations?
- Do higher asymptote recipes unlock genuinely novel reasoning strategies?
- Why do reasoning models fail to improve constrained optimization performance?
- How does soft thinking compare to sampling multiple independent reasoning paths?
- How do soft token mixtures enable parallel reasoning exploration without explicit training?
- How does soft thinking achieve stochastic exploration without explicit training?
- Can one training example activate mathematical reasoning without reinforcement learning?
- How can verifier-free reinforcement learning handle reasoning without task-specific checks?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can inference compute replace scaling up model size?
Explores whether smaller models given more thinking time during inference can match larger models. Matters because it reshapes deployment economics and compute allocation strategies.
EBTs operationalize this at the architecture level: energy minimization inherently scales with inference compute
-
Can non-reasoning models catch up with more compute?
Explores whether inference-time compute budget can close the performance gap between standard models and those trained for reasoning, and what training mechanisms might enable this.
EBTs may redefine the boundary: energy minimization is a form of inference-time computation that doesn't require reasoning-specific RL training
-
Does more thinking time actually improve LLM reasoning?
The intuition that extended thinking helps LLMs reason better seems obvious, but what does the empirical data actually show when we test it directly?
EBTs add nuance: for energy-based architectures, more iterations genuinely improve until convergence, unlike token-based reasoning where overthinking degrades quality
-
Can recurrent hierarchies achieve reasoning that transformers cannot?
Can a dual-timescale recurrent architecture escape the computational limitations of standard transformers and solve complex reasoning tasks without explicit chain-of-thought? This explores whether architectural design, not scale, enables true algorithmic reasoning.
complementary latent architecture: HRM achieves near-perfect accuracy on tasks where CoT scores 0% via dual-recurrence; EBTs achieve 35% higher scaling rate via energy minimization; different mechanisms (recurrence vs. gradient descent) escaping the same TC0 constraint
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Energy-Based Transformers are Scalable Learners and Thinkers
- Grokked Transformers are Implicit Reasoners: A Mechanistic Journey to the Edge of Generalization
- Hierarchical Reasoning Model
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
- Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory
- Can Large Language Models Reason and Optimize Under Constraints?
- Farther the Shift, Sparser the Representation: Analyzing OOD Mechanisms in LLMs
- Performative Thinking? The Brittle Correlation Between CoT Length and Problem Complexity
Original note title
energy-based transformers achieve system 2 thinking from unsupervised learning alone — modality and problem agnostic