SYNTHESIS NOTE
Topics›Alignment›this note

Can models learn to ignore irrelevant prompt changes?

Explores whether training models to produce consistent outputs regardless of sycophantic cues or jailbreak wrappers can solve alignment problems rooted in attention bias rather than capability gaps.

Synthesis note · 2026-02-23 · sourced from Alignment

Sycophancy and jailbreaking share a structural property: the model produces the correct response to a clean prompt but changes its response when irrelevant cues are added (a user's stated opinion, a jailbreak wrapper). The problem is not capability — it's consistency.

Consistency training reframes alignment as invariance: train the model to produce the same response regardless of whether the prompt includes irrelevant perturbations. Two methods implement this:

Bias-Augmented Consistency Training (BCT) operates on output tokens. For each prompt, the model generates a response to the clean version. This response becomes the training target for the wrapped version. The model learns to say the same thing regardless of sycophantic cues.

Activation Consistency Training (ACT) operates on internal representations. Instead of matching output tokens, ACT enforces that residual stream activations on the wrapped prompt match those on the clean prompt. This is a more mechanistic constraint — teaching the model to think the same way, not just say the same thing.

Both reduce sycophancy effectively. BCT is better at jailbreak reduction. The advantage over standard SFT is avoiding two forms of staleness:

Since consistency training uses the model's own clean responses as targets, both staleness problems disappear. The training data is always fresh and at the model's current capability level.

Continual learning extension — Self-Distillation Fine-Tuning (SDFT). SDFT generalizes the self-as-target principle to continual learning from demonstrations. The model plays two roles: a teacher conditioned on both input and expert demonstration (via in-context learning), and a student conditioned on input only. Training distills the teacher into the student on trajectories generated by the student itself — yielding on-policy updates that incorporate demonstration knowledge without explicit reward inference. SDFT achieves higher new-task accuracy while substantially reducing catastrophic forgetting vs standard SFT. In sequential learning across three skills, a single model accumulates each skill without regression on previously learned abilities. The mechanism parallels BCT: both use the model's own contextually-enhanced output as the training signal, avoiding off-policy distribution mismatch.

This connects to Does transformer attention architecture inherently favor repeated content?. S2A identifies the architectural root (attention bias toward repeated/prominent tokens); consistency training provides the training-level fix (enforce invariance to those biased attention patterns). ACT's activation-level approach is particularly relevant — it may directly counteract the attention bias at the representation level.

ProSA (2024) provides the diagnostic that explains WHY consistency training works. Prompt sensitivity is fundamentally a reflection of model confidence: higher confidence correlates with increased robustness against prompt semantic variations. This means consistency training (BCT/ACT) succeeds not by teaching a separate "invariance skill" but by pushing models toward confident response regions where robustness is a natural property. Few-shot examples also alleviate sensitivity by providing concrete anchoring. Larger models exhibit enhanced robustness. Source: Arxiv/Prompts Prompting.

Inquiring lines that read this note 147

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do prompting refinements mask underlying biases and model frequency patterns? Can prompt-based context override biases that were embedded during pretraining? How do surface patterns enable correct outputs but reduce robustness? How do evaluation practices shape which failures stay visible? Can self-generated feedback reliably guide model training without ground truth? Why do LLM recommenders underperform collaborative filtering despite their capabilities? Does transformer attention architecture inherently drive sycophancy? How do false presuppositions and sycophancy drive persistent false beliefs in models? Can compression size predict model complexity better than parameter count alone? Why do stronger reasoning capabilities create tradeoffs with instruction following? Do language models learn genuine understanding or just surface patterns? How do LLM judges' systematic biases affect alignment and evaluation outcomes? What emerges when safety-aligned models attempt to role-play deceptive personas? Why can't prompting alone inject genuinely new knowledge into models? What training dynamics and scale trigger emergence of reasoning capabilities? What determines appropriate intervention timing and manner for AI agents? How do prompt design choices influence model reasoning and performance? What structural properties of attention create systematic model biases? When do semantic similarity approaches miss structural retrieval failures? How do training data properties determine the emergence of internal misalignment? Can inoculation prompting prevent emergent misalignment after reward hacking? Can reasoning scale in latent space without tokens? How can conversational agents maintain consistent personas across multi-turn dialogue? Does alignment training create genuine alignment or just output compliance? What attack surfaces do reasoning traces and chains introduce? Can mechanistic interpretability reliably guide practical model design choices? How can reward models capture diverse human preferences without excluding minority populations? Does preference optimization systematically degrade conversational grounding in language models? How does persona conditioning amplify demographic stereotyping and bias in models? How do spurious versus genuine rewards shape model reasoning and behavior? What mechanisms preserve shared understanding in evolving conversations? Why don't LLMs reliably translate capability into accurate outputs? How can we distinguish genuine model deception from honest errors? Does RLHF training systematically drive models toward sycophancy and away from accuracy? How can oversight detect and prevent conditional compliance when agents know they are watched? What makes distillation transfer some model capabilities while suppressing others? Can preference-based training achieve better behavior optimization than supervised fine-tuning alone? Do reasoning benchmarks predict model performance in long-horizon workflows? How does misalignment propagate through agent communication networks? How can we build reliable evaluations of AI reasoning despite judge bias and reward-seeking? What trajectory-level metrics beyond task success best evaluate agent performance? Do reasoning traces faithfully reflect actual model reasoning? Does encoded knowledge in language models actually influence their outputs? Do language models respond to social pressure and face-saving like humans?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
17 direct connections · 186 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

consistency training teaches models prompt-perturbation invariance using their own clean responses as targets — avoiding SFT staleness