SYNTHESIS NOTE
Topics›Training Fine Tuning›this note

Can utility-weighted training loss actually harm model performance?

When engineers weight loss functions to reflect real-world costs of different errors, does this improve or undermine learning? This explores whether baking asymmetric objectives into training creates unintended side effects.

Synthesis note · 2026-02-22 · sourced from Training Fine Tuning

"Misaligned by Design" identifies a failure in what the authors call the Aligned Learning Premise (ALP): the intuition that using the human's utility function to train a model produces better performance in terms of that objective. In high-stakes settings where false positives and false negatives have asymmetric costs (e.g., medical diagnosis), engineers routinely bake these asymmetric weights into the training loss. This paper shows this can backfire.

The key insight: machine classifiers perform not one but two incentivized tasks. Choosing how to classify (given learned features, assign a label) — here asymmetric weighting works correctly. Learning how to classify (acquiring informative feature representations through gradient descent) — here asymmetric weighting can weaken the learning signal. Because the loss function shapes the gradient, it necessarily shapes incentives for learning. Making the loss asymmetric can reduce the payoff to "substantive learning" — the model learns less informative representations.

In both focal applications, training with a standard symmetric loss function then adjusting predictions ex-post according to the human's utility function outperforms training with the utility-weighted loss directly — even when evaluated by the utility-weighted objective itself. Trying to bake utility weights into training makes predictions worse.

This resonates with findings across the LLM training literature. Do reward models actually consider what the prompt asks? shows reward models that should evaluate answer quality actually ignore the question — an incentive misalignment between what the loss teaches and what the evaluation requires. Does supervised fine-tuning actually improve reasoning quality? shows SFT optimizing for accuracy inadvertently degrades reasoning quality — the loss correctly incentivizes choosing the right answer but weakens the incentive to learn informative reasoning paths.

The general principle: when a training objective conflates two functions (learning representations and making decisions), optimizing one can degrade the other. Separating them — learn first, then decide — may be superior even though it seems less elegant. The Does binary reward training hurt model calibration? finding is a direct instance: binary reward correctly incentivizes choosing (pick the right answer) but fails to incentivize learning calibrated confidence, and the Brier score fix explicitly separates these two objectives within the reward function.

Inquiring lines that read this note 27

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Do structural constraints outperform deep architectures in recommendation systems? How do neural networks achieve compositional generalization at scale? How do pretraining biases affect reward signal effectiveness in RLVR? Can preference-based training achieve better behavior optimization than supervised fine-tuning alone? How do recommenders balance exploiting fresh signals against maintaining preference stability? Does model confidence reliably signal actual accuracy in practice? How do capability benchmark scores systematically misrepresent true model abilities? What training data selection strategies maximize generalization across difficulty levels? Does alignment training create genuine alignment or just output compliance? Can iterative DPO replicate online reinforcement learning dynamics for research? Does warmth and empathy training systematically degrade model reliability? How does harness optimization generalize across different model architectures and domains? How do surface patterns enable correct outputs but reduce robustness? What capability trade-offs arise from domain specialization through fine-tuning? Why do token-level mechanisms matter for learning to reason? What training dynamics and scale trigger emergence of reasoning capabilities?

Related concepts in this collection 7

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
15 direct connections · 198 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

asymmetric loss functions can misalign machine learning because learning and choosing are distinct incentivized tasks — utility-weighted training can backfire