Can scalar rewards capture all the information in agent feedback?
Exploring whether numerical rewards alone can preserve both the evaluative judgment and directional guidance embedded in natural feedback—or if something crucial gets lost in the conversion.
The OpenClaw-RL framework makes a decomposition that was implicit in prior agentic RL work but never formalized: when an agent acts and the environment responds, the response carries two distinct kinds of information. The evaluative signal scores the action — how well did it perform — and can be extracted as a scalar reward via a PRM judge. The directive signal specifies how the action should have been different — not just that it was wrong, but in what direction. These are orthogonal: high-quality directive information can accompany any evaluation, and scalar rewards systematically lose the directive component.
Consider a user who says "you should have checked the file first." The evaluative content is approximately -1 (the response was inadequate). But the directive content is token-level specific: check the file first. A PRM judge can convert the sentiment into a scalar, but the sequence-level correction vanishes into a single number. Similarly, a detailed SWE error trace often implies a concrete correction direction that scalar outcome rewards cannot convey. Current RLVR methods operate on scalar rewards (Does RLVR actually expand what models can reason about?) and cannot convert directive information into a directional policy gradient. Distillation methods can process structured corrections but require pre-curated feedback-response pairs rather than live signals.
OpenClaw-RL recovers the directive signal through Hindsight-Guided On-Policy Distillation (OPD): extract textual hints from the next state, construct an enhanced teacher context by injecting those hints, and distill token-level directional advantage back into the student policy. This is richer than any scalar reward because it teaches the model not just "that was wrong" but "here is what right looks like in these specific tokens." The empirical result — combining binary PRM-based RL with OPD via weighted loss yields significant gains over either alone — confirms the two signals are complementary, not redundant.
This decomposition matters beyond OpenClaw-RL because it clarifies a conceptual muddle in agentic RL. When people debate "should we use outcome rewards or process rewards, scalar or verbal," the answer is usually "both, decomposed properly." The outcome-vs-process trade-off (Why do outcome-based reward models fail at intermediate step evaluation?) assumes a single signal type. The scalar-vs-verbal distinction is treated as architectural (Can natural language feedback overcome numerical reward plateaus?). OpenClaw-RL reframes them as two projections of one signal: evaluative (dense scalar) and directive (token-level).
The generalization: any learning loop that reduces natural feedback to scalars is discarding the fraction of training signal that most resembles supervised learning. A corrective sentence contains its own teacher.
Inquiring lines that read this note 189
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do neighboring agents influence whether others cooperate or collude?- Do explicit reward structures enable AI agent cooperation that open-ended interaction cannot?
- Does social scaffolding outperform purely intrinsic motivation for agent exploration?
- How does asymmetric information between users and agents relate to proactivity?
- Can agents learn cooperation from reward signals alone without seeing helpers?
- Why do counterfactual credit methods fail on unobserved cooperation?
- What cognitive capabilities do agents need to internalize social feedback?
- What makes exploration and reflection rewards verifiable in agentic environments?
- Why do scalar evaluation scores collapse distinguishable agent behaviors?
- How do agent actions change state that reward procedures later read?
- Why does moving the reward target prevent saturation better than finding a better static proxy?
- Can an agent change reward-path state through actions during evaluation?
- Should feedback channels be excluded from the reward path in agent evaluations?
- Can unified policies handle negative feedback and critique transformation simultaneously?
- Can systems recognize and abstain on judgments rather than hallucinating preferences?
- How does information asymmetry between teacher and student create the learning signal?
- Why does information asymmetry between teacher and student enable effective feedback learning?
- Can environment feedback alone provide dense credit without a teacher?
- How does credit assignment drive agents to write information into environments?
- Do agents prefer raw experience over condensed summaries of past actions?
- How can agents distinguish over-generalized lessons from genuinely useful long-tail knowledge?
- Why does binary reward forcing degrade model calibration?
- How does RLHF reward structure incentivize agreement over accuracy?
- Can distillation methods extract directional guidance that scalar RL cannot access?
- Does in-distribution reward model performance hide failures from context shift?
- How do reward model ensembles improve robustness to miscalibration?
- Can reward engineering and information-theoretic architecture solve partner-awareness separately?
- Can reward model training be automated without changing feedback mechanisms?
- What information do next-state signals contain beyond what scalar rewards capture?
- How does modularity in reward and policy design enable goal generalization?
- At what capability level does the generation-verification gap make intrinsic rewards insufficient?
- What distinguishes verifiable rewards from preference-based rewards in unified training?
- How does negative reinforcement redistribute probability without guiding toward correct answers?
- Is elaborate reward shaping necessary if the pretrained prior already contains good solutions?
- Can model confidence signals replace explicit external reward functions?
- How do reward model biases cascade into downstream optimization failures?
- Why do next-turn reward objectives fail to encourage multi-turn goal progress?
- Why do spurious rewards work nearly as well as correct ones?
- What makes Effective Rank Acceleration a stable training signal for dual-channel incentives?
- Can reward factorization represent trade-offs between conflicting moral values?
- What reward mechanisms make thinking-based compression budget-controllable and reliable?
- How does belief-shift reward compare to curiosity-driven and process reward approaches?
- Can an agent's internal probabilities serve as value signals across domains?
- Why does scalarization of rewards fail for multi-objective GRPO training?
- What other downstream metrics could serve as RL reward sources?
- How do you extract reward signals when all rollouts fail?
- Can verifiable rewards during pretraining replace costly human preference labeling?
- How do relational reward signals compare to absolute preference encodings in RL?
- How does in-context feedback integration differ from learned reward signals?
- Can early experience replace external rewards as a learning signal?
- What other adaptive internal phenomena could signal system behavior improvements?
- Can tree-GRPO work with extremely noisy or sparse outcome reward signals?
- What makes reward models fundamentally different from policy discriminators?
- How does DVAO balance reward components differently than VPO spreads them?
- When does a task lack a meaningful multi-dimensional reward structure?
- What makes reward signal sources substitutable across verifier-free RL patterns?
- What makes advantage shaping more stable than reward shaping for tool training?
- What makes a sub-goal verifiable enough to provide dense feedback signals?
- How do reward models and self-improvement mechanisms interact in training?
- Why do outcome-only rewards fail to optimize long-horizon agent behavior?
- Why do dense rewards plus hard constraints outperform single fixed rewards?
- How do self-play and human-anchored rewards separate competence from convention?
- Does inverse-variance denoising reduce variance below either reward stream alone?
- What mechanisms do peer predictions use to generate reward signals for training?
- Can simple intrinsic reward signals emerge as effective drivers of complex capability in agents?
- Do spurious rewards activate reasoning without teaching new skills?
- Can multi-turn rewards fix models that lose track midway?
- Why does natural language feedback break performance plateaus that numerical rewards alone cannot?
- Do outcome-only reward signals miss step-level errors that compound later?
- How do graduated phase rewards emerge complex dialogue behavior from simple objectives?
- How does reward model training permit spurious correlations in scoring?
- Can reward models trained for engagement fix the informativeness problem?
- How do semantic reward shaping approaches compare to full critique models?
- What information do numerical rewards fail to provide for reasoning tasks?
- Why do generative reward models produce more interpretable evaluations than scalar scores?
- Why do reward models fail when they ignore the prompt context?
- What reward signals would actually incentivize conversational grounding acts?
- How can reward structures teach models when to speak and when to stay silent?
- How do task-type perceptions like chat versus reasoning guide different reward strategies?
- How do reward models benefit from extended thinking during evaluation scoring?
- Can structured natural language feedback outperform scalar rewards in RL?
- Can multi-turn aware rewards improve alignment beyond single-turn helpfulness?
- Can decomposing rewards into prompt-free and prompt-related components fix this blindspot?
- Can reward design fix the conflict between reasoning accuracy and abstention calibration?
- How does reinforcement learning on outcomes reinforce template-matching rather than computation?
- Can environmental rewards directly refine natural language descriptions of actions?
- How do checklists prevent reward models from exploiting superficial response artifacts?
- How do token-level rewards and rubric gates serve different statistical functions?
- Can structured rewards still teach models when spurious rewards also work?
- What role does task structure play in rewarding delayed thinking?
- What makes binary rewards more effective than richer reward signals?
- What makes user-decision rewards better than model-confidence rewards?
- Do information gathering and task execution require different incentive structures?
- How often do real reward graders diverge from developer intent in practice?
- Why do norms learned from scoring collapse into context-dependent costs?
- Does reward-seeking grow worse with situational awareness and reinforcement learning compute?
- Does length bias in reward models explain response growth across iterations?
- What makes an agent notice that reward beats compliance?
- Why do weak belief tracking and conservative actions trap agents in low-information states?
- Why do completion-mode strengths not transfer to agentic settings?
- Can agents escape weak belief tracking and conservative action selection traps?
- How do outcome and process rewards differ in their treatment of intermediate steps?
- Can solution traces substitute for process-level reward signals in math reasoning?
- What makes process-level supervision better than outcome-only reward signals?
- How does process-focused feedback compare to outcome-focused feedback in skill training?
- How do process-level rewards compare to environment-extracted next-state signals?
- Can programmatic meta-reasoning rewards operationalize agentic process supervision?
- What information-theoretic framework explains why process rewards beat outcome only?
- What distinguishes generative reward models from outcome-based and process-based approaches?
- How do outcome-based and process-based reward models differ in supervision cost?
- How does tree-search topology convert outcome rewards into intermediate supervision?
- How does belief-shift credit assignment compare to process reward models?
- Do process reward models need different supervision strategies by domain?
- Can trajectory structure replace hand-annotated process reward models entirely?
- How does process-based reward differ from outcome-only reward in training?
- How can consistency across measurement conditions identify genuine versus constructed preferences?
- Can negative feedback through critiques achieve the same steering flexibility as positive preferences?
- Can rich environment feedback replace human preference labels entirely?
- Can light human signals steer already-learned behavior without preference labels?
- Can importance sampling reduce variance in off-policy reward estimation?
- Can curiosity rewards about user type complement general social motivation frameworks?
- What preference dimensions do base reward functions typically capture?
- Can we distinguish between genuine alignment and response quality bias in reward signals?
- How do reward features learned from group data generalize to new users?
- How do reward models as policy discriminators differ from labeled preferences?
- Can vector-valued rewards preserve specialization better than variance-weighted advantages?
- Can user preferences be represented as linear reward combinations?
- Can reward models distinguish between personal preference and community consensus?
- Do personalized reward models work better than one-size-fits-all approaches?
- Does pairwise self-judgment avoid reward model scaling problems?
- How do aggregate reward models systematically exclude minority perspectives?
- How do aggregate reward models systematically exclude minority preferences?
- What makes trajectory more actionable than absolute scores for human moderators?
- What deployment modes work best for trajectory-aware reward signals?
- Why do sparse outcome rewards fail to credit correct tool calls in failed trajectories?
- What hidden signals in agent logs reveal about frontier capability beyond pass-fail outcomes?
- Should agent evaluation include trajectory quality beyond final success?
- Does decision-making taste predict end-to-end task success independently?
- What are the ten intrinsic motivation heuristics that drive participation decisions?
- Can evasive non-commitment mask withheld feedback while appearing thoughtful?
- Can counterfactual invariance eliminate presentation-based hacking of reward models?
- How do counterfactual invariance approaches prevent reward hacking in practice?
- Why do veto mechanisms on critical dimensions prevent collapse into exploitable reward modes?
- Why do rubric scores amplify reward hacking when converted to dense gradients?
- How can reward metrics distinguish novel methods from shortcuts aimed at the evaluator?
- How does implicit feedback structure differ from explicit ratings mathematically?
- How do confidence signals differ between implicit feedback and explicit ratings?
- How do evaluative versus directive signals differ in next-state training?
- How do weights, selection, and prompts create different geometric landscapes of accessible behaviors?
- Do fed-back concepts or the auxiliary objective alone drive the performance gain?
- Can architectural changes like decoupling intent understanding help overcome next-turn reward limitations?
- What design changes if we separate behavior description from adoption justification goals?
- Can agents revise their beliefs predictably when presented with interventions?
- How does next-turn reward optimization contribute to agent passivity?
- How do human-agent systems incorporate diverse feedback into model behavior?
- How do you prevent stale reward signals when skills evolve during deployment?
- How does effective feedback retention govern long-horizon agent reliability?
- Why does persistence in the feedback loop predict agent success better than initial solution quality?
- Why do agents fail to internalize value from informative observations?
- How do delayed effects complicate causal attribution in agent systems?
- How does credit assignment across objectives differ from credit assignment across time?
- What separates bootstrapping gains from sustained self-improvement gains?
- Does the generation-verification gap define where self-rewarding actually works?
- How does Goodhart's Law apply to proxy rewards in self-training systems?
- Why does research-direction judgment validation limit fully closed self-improvement?
- Does self-play feedback improve skills created from the agent's own experience?
- Why do self-improving agents concentrate progress in the fast non-parametric loop?
- Does fixed evaluation criteria saturate as self-improving agents improve?
- How should multi-objective post-training balance competing behavioral goals?
- What alignment properties emerge when the reward model disappears?
Related concepts in this collection 6
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can agent deployment itself generate training signals automatically?
Can we extract learning signals from the natural next-states that agents encounter during real deployment—user replies, tool outputs, test verdicts—rather than relying on separate annotation pipelines? This reframes how agents improve continuously.
the framing this decomposition operates within
-
Can natural language feedback overcome numerical reward plateaus?
Exploring whether chain-of-thought critiques can push past performance ceilings that scaling data alone cannot break in reinforcement learning for reasoning tasks.
establishes that verbal feedback contains information scalars cannot reach
-
Does binary reward training hurt model calibration?
Explores whether the standard correctness-based reward in RL training creates incentives for overconfident predictions, and what structural problem causes calibration to degrade during optimization.
another case where single-scalar objectives miss structure
-
Why do outcome-based reward models fail at intermediate step evaluation?
Outcome-based reward models (ORMs) evaluate only final results, creating a mismatch with the need to assess reasoning quality at intermediate steps. Understanding this failure mode matters for building better AI reasoning systems.
the outcome/process axis is the wrong cut; evaluative/directive is closer to the information structure
-
Does critiquing errors teach deeper understanding than imitating correct answers?
Can training models to critique flawed responses build better structural understanding than standard supervised fine-tuning on correct answers? This matters because it reveals whether deep reasoning requires engaging with failure modes rather than pattern matching.
critique-based training as a cousin: teaching the model the directive structure behind errors
-
Does RLVR actually expand what models can reason about?
Explores whether reinforcement learning from verifiable rewards teaches models genuinely new reasoning skills or simply makes existing capabilities more reliable. Pass@k analysis suggests the latter.
scalar RLVR's structural ceiling that directive signals may penetrate
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Reward Reasoning Model
- OpenClaw-RL: Train Any Agent Simply by Talking
- A Survey of Reinforcement Learning from Human Feedback
- Reinforcement Learning via Self-Distillation
- Information-Theoretic Reward Decomposition for Generalizable RLHF
- Can Large Language Models Reason and Optimize Under Constraints?
- Foundations of Large Language Models
- Online Intrinsic Rewards for Decision Making Agents from Large Language Model Feedback
Original note title
agent next-state signals decompose into evaluative and directive information that scalar rewards cannot jointly capture