SYNTHESIS NOTE
Topics›Reasoning by Reflection›this note

Can confidence patterns reveal overthinking versus underthinking?

This explores whether real-time confidence signals can diagnose when a reasoning model is trapped in redundant deliberation versus committing prematurely, and whether steering based on these signals can balance both failure modes.

Synthesis note · 2026-04-01 · sourced from Reasoning by Reflection

Overthinking and underthinking are dual failures, and existing methods that suppress one often induce the other. Suppressing reflective keywords or truncating reasoning length reduces overthinking but causes underthinking — the model doesn't explore enough. Forcing longer chains reduces underthinking but generates redundancy. ReBalance resolves this by treating confidence as a continuous diagnostic signal rather than using binary interventions.

The diagnostic: Confidence values correlate with reasoning behavior in interpretable ways:

The mechanism: From a small-scale dataset, identify reasoning steps indicating each mode. Aggregate their hidden states into reasoning mode prototypes. Compute a steering vector encoding the transition from overthinking to underthinking. A dynamic control function modulates the vector's strength and direction based on real-time confidence: pruning redundancy during overthinking, promoting exploration during underthinking.

Why it's training-free: The steering vector captures the model's inherent reasoning dynamics — it's extracted from the model's own hidden states, not trained. Because it operates on intrinsic representations, it generalizes across unseen data and tasks (math, QA, coding). This makes it plug-and-play across models from 0.5B to 32B.

Since Can we steer reasoning toward brevity without retraining?, ReBalance extends the activation-steering approach from length compression to reasoning quality management. ASC steers between verbose and concise modes; ReBalance steers between overthinking and underthinking — a qualitative distinction, not just quantitative.

Since Does more thinking time always improve reasoning accuracy?, ReBalance provides the dynamic mechanism the threshold finding calls for: instead of a fixed cutoff, confidence-based steering continuously adjusts the reasoning trajectory.

Inquiring lines that read this note 85

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do false presuppositions and sycophancy drive persistent false beliefs in models? How does reasoning length affect model performance across different tasks? Does chain-of-thought reasoning reveal genuine computation or imitate patterns? Can harness architecture and protocols provide agent reliability without model scaling? What determines appropriate intervention timing and manner for AI agents? Does model confidence reliably signal actual accuracy in practice? What reasoning architectures enable models to solve complex problems efficiently? Why do some clarifying approaches produce understanding while others just satisfy? Why do token-level mechanisms matter for learning to reason? What is the relationship between thinking tokens and reasoning accuracy? Do reasoning traces faithfully reflect actual model reasoning? How should systems decide whether to retrieve or reason alone? How does self-revision in reasoning models affect accuracy and confidence? Does preference optimization systematically degrade conversational grounding in language models? What structural distinctions matter in reasoning and argumentation? Is reasoning capability latent in base models or created by post-training? Why do persona simulations fail to predict authentic user behavior? What causes reasoning models to fail or wander off track? Why do agents falsely report success on failed tasks? Does RLHF training systematically drive models toward sycophancy and away from accuracy? Do language models reason like humans or mimic surface patterns? Can brute-force automated research substitute for iterative depth and human research intuition? Can self-generated feedback reliably guide model training without ground truth? How should designers communicate what AI systems truly are and can do? How does improved reasoning affect models' ability to acknowledge uncertainty? Why does polished presentation create unearned authority in AI outputs? Can models improve accuracy without degrading reasoning quality? What makes distillation transfer some model capabilities while suppressing others? How do soft reasoning mechanisms explore multiple paths without explicit training? Does RL create genuinely new reasoning capabilities or refine existing ones? Can inference-time compute effectively substitute for model scale? Can multi-agent systems avoid converging on false agreement without deliberation? Can reasoning traces and behavior monitoring reliably detect hidden AI scheming? What trajectory-level metrics beyond task success best evaluate agent performance?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 103 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

ReBalance uses confidence as continuous indicator to dynamically steer between overthinking and underthinking — training-free balanced reasoning via hidden state steering vectors