SYNTHESIS NOTE
Topics›Autonomous Agents›this note

Can scalar rewards capture all the information in agent feedback?

Exploring whether numerical rewards alone can preserve both the evaluative judgment and directional guidance embedded in natural feedback—or if something crucial gets lost in the conversion.

Synthesis note · 2026-04-07 · sourced from Autonomous Agents

The OpenClaw-RL framework makes a decomposition that was implicit in prior agentic RL work but never formalized: when an agent acts and the environment responds, the response carries two distinct kinds of information. The evaluative signal scores the action — how well did it perform — and can be extracted as a scalar reward via a PRM judge. The directive signal specifies how the action should have been different — not just that it was wrong, but in what direction. These are orthogonal: high-quality directive information can accompany any evaluation, and scalar rewards systematically lose the directive component.

Consider a user who says "you should have checked the file first." The evaluative content is approximately -1 (the response was inadequate). But the directive content is token-level specific: check the file first. A PRM judge can convert the sentiment into a scalar, but the sequence-level correction vanishes into a single number. Similarly, a detailed SWE error trace often implies a concrete correction direction that scalar outcome rewards cannot convey. Current RLVR methods operate on scalar rewards (Does RLVR actually expand what models can reason about?) and cannot convert directive information into a directional policy gradient. Distillation methods can process structured corrections but require pre-curated feedback-response pairs rather than live signals.

OpenClaw-RL recovers the directive signal through Hindsight-Guided On-Policy Distillation (OPD): extract textual hints from the next state, construct an enhanced teacher context by injecting those hints, and distill token-level directional advantage back into the student policy. This is richer than any scalar reward because it teaches the model not just "that was wrong" but "here is what right looks like in these specific tokens." The empirical result — combining binary PRM-based RL with OPD via weighted loss yields significant gains over either alone — confirms the two signals are complementary, not redundant.

This decomposition matters beyond OpenClaw-RL because it clarifies a conceptual muddle in agentic RL. When people debate "should we use outcome rewards or process rewards, scalar or verbal," the answer is usually "both, decomposed properly." The outcome-vs-process trade-off (Why do outcome-based reward models fail at intermediate step evaluation?) assumes a single signal type. The scalar-vs-verbal distinction is treated as architectural (Can natural language feedback overcome numerical reward plateaus?). OpenClaw-RL reframes them as two projections of one signal: evaluative (dense scalar) and directive (token-level).

The generalization: any learning loop that reduces natural feedback to scalars is discarding the fraction of training signal that most resembles supervised learning. A corrective sentence contains its own teacher.

Inquiring lines that read this note 189

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do neighboring agents influence whether others cooperate or collude? What should agent evaluation prioritize to reveal reliable behavior? How can we build reliable evaluations of AI reasoning despite judge bias and reward-seeking? How does evaluation scope and dimensionality affect what we measure? How should agents manage memory granularity to improve long-term performance? How do pretraining biases affect reward signal effectiveness in RLVR? How do spurious versus genuine rewards shape model reasoning and behavior? Can harness architecture and protocols provide agent reliability without model scaling? What makes step-level supervision effective for complex reasoning traces? Can preference-based training achieve better behavior optimization than supervised fine-tuning alone? Why do people disclose to AI systems despite their artificial nature? How can reward models capture diverse human preferences without excluding minority populations? What trajectory-level metrics beyond task success best evaluate agent performance? How do recommenders balance exploiting fresh signals against maintaining preference stability? What determines appropriate intervention timing and manner for AI agents? How can humans maintain meaningful oversight as AI systems become increasingly autonomous and complex? Can we reliably detect when models game evaluations? How do social dynamics distort aggregated online ratings? What training dynamics and scale trigger emergence of reasoning capabilities? Should agents decouple planning from perception grounding for better performance? How does policy entropy collapse constrain scaling of reasoning-focused RL? What makes personas effective for predicting individual preferences and behavior? How do agent-learned skills transfer and improve across different tasks? Why do agents falsely report success on failed tasks? Do language models reason like humans or mimic surface patterns? Why does memory consolidation cause performance regression in continual learning? Can self-generated feedback reliably guide model training without ground truth? Does warmth and empathy training systematically degrade model reliability? How should agent systems validate and persist generated code artifacts? What fundamental constraints limit how effectively agents can improve themselves? Does alignment training create genuine alignment or just output compliance? How do surface patterns enable correct outputs but reduce robustness? What do systematic disagreements between annotators reveal about ground truth? How does the generation-verification gap limit what we can measure about AI reasoning? How can infrastructure records verify actual agent behavior? Why do standard benchmarks fail to predict agent deployment success? Can welfare maximization and minority veto protection coexist? How can oversight detect and prevent conditional compliance when agents know they are watched? Can multi-agent systems avoid converging on false agreement without deliberation? Does AI assistance promote real skill development or substitute for independent learning?

Related concepts in this collection 6

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
18 direct connections · 157 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

agent next-state signals decompose into evaluative and directive information that scalar rewards cannot jointly capture