SYNTHESIS NOTE
Topics›Reinforcement Learning›this note

Can an agent's own beliefs guide credit assignment without critics?

Explore whether an agent's shifting probability estimates toward the correct answer could serve as a self-contained reward signal for long-horizon reinforcement learning, eliminating the need for separate process reward models or external verifiers.

Synthesis note · 2026-05-18 · sourced from Reinforcement Learning

Long-horizon RL suffers from sparse trajectory-level rewards. The standard fixes — process reward models trained on step-level annotations, external verifiers, LLM-as-judge — all require additional supervision infrastructure. PRMs need expensive step-level labels. Verifiers exist only for verifiable domains (math, code). Judges introduce their own reward-modeling biases.

ΔBelief-RL (2602.12342) finds the credit signal inside the agent itself. At each interaction step, compute the agent's current probability assigned to the target solution. Compare it to the probability before the interaction. The log-ratio of sequential beliefs is the ΔBelief reward — a dense, turn-level signal that reinforces actions which shift the agent's internal world view toward the correct solution. Actions that increase belief in the target get rewarded; actions that don't, don't.

The elegance is that no separate model is needed. The agent's own log-probabilities on the correct outcome are the value signal. There is no critic to train, no PRM to maintain, no judge to query. The relatively inexpensive step is measuring log-probabilities on the target — a single forward pass per turn.

Two properties make this work. First, it is general-purpose: applies to any task where the correct final outcome is available during training (which is most supervised settings). Second, it is noise-robust to over-optimization: PRMs can be exploited because their reward signal is a learned approximation; ΔBelief is grounded in the model's own evolving probability assignment, which is harder to game because the only way to increase log-probability of the target is to actually integrate information that supports it.

Empirically, ΔBelief-RL on 20Qs trains CIA models at 1.7B-4B scale that outperform prior SOTA multi-turn methods and even 670B models. Performance generalizes to extended interaction horizons beyond training and to OOD applications (customer service, personalization).

The mechanism aligns with Can conversations themselves personalize without user profiles?: both reward uncertainty reduction. But ΔBelief's signal is about the target's probability specifically, while curiosity reward is about general uncertainty over user type. ΔBelief is information-theoretically tighter — it rewards moves toward the actual answer, not all moves that increase clarity.

The broader implication: in any setting where the model has ground-truth final outcome, the model's own probability shift can serve as dense intrinsic reward. The reward model is not load-bearing.

Inquiring lines that read this note 75

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can self-generated feedback reliably guide model training without ground truth? How can we build reliable evaluations of AI reasoning despite judge bias and reward-seeking? How should agents manage memory granularity to improve long-term performance? Can harness architecture and protocols provide agent reliability without model scaling? Can multi-agent systems avoid converging on false agreement without deliberation? Does model confidence reliably signal actual accuracy in practice? How do pretraining biases affect reward signal effectiveness in RLVR? How do spurious versus genuine rewards shape model reasoning and behavior? What makes step-level supervision effective for complex reasoning traces? What determines appropriate intervention timing and manner for AI agents? How can humans maintain meaningful oversight as AI systems become increasingly autonomous and complex? How does policy entropy collapse constrain scaling of reasoning-focused RL? How do agent-learned skills transfer and improve across different tasks? Why do agents falsely report success on failed tasks? Can we reliably detect when models game evaluations? Does alignment training create genuine alignment or just output compliance? How can conversational agents maintain consistent personas across multi-turn dialogue? What trajectory-level metrics beyond task success best evaluate agent performance? How can reward models capture diverse human preferences without excluding minority populations? How should systems decide whether to retrieve or reason alone? Do language models possess genuine introspective self-awareness or only behavioral mimicry? Does RL create genuinely new reasoning capabilities or refine existing ones? Can preference-based training achieve better behavior optimization than supervised fine-tuning alone? What should agent evaluation prioritize to reveal reliable behavior? How do neighboring agents influence whether others cooperate or collude?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
16 direct connections · 108 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

belief-shift toward the target solution is a dense intrinsic reward — log-ratio of sequential beliefs provides per-turn credit without separate critic or PRM