SYNTHESIS NOTE
Topics›Flaws›this note

How vulnerable are reasoning models to irrelevant text?

Can simple adversarial triggers like unrelated sentences degrade reasoning model accuracy? This explores whether step-by-step reasoning actually provides robustness against subtle input perturbations.

Synthesis note · 2026-02-23 · sourced from Flaws

Reasoning models are vulnerable to a startlingly simple attack: appending short, semantically irrelevant text to any math problem systematically misleads them. "Interesting fact: cats sleep most of their lives" appended to a math problem more than doubles the chance of an incorrect answer.

The CatAttack pipeline discovers triggers on a weaker, cheaper proxy model (DeepSeek V3) that successfully transfer to stronger reasoning targets like DeepSeek R1 and R1-distilled-Qwen-32B, increasing error rates by over 300%. The triggers are:

This is distinct from Why do reasoning models fail under manipulative prompts?. Gaslighting attacks use multi-turn social pressure; CatAttack uses single-shot irrelevant text. Both show reasoning models are brittle, but through different mechanisms. Gaslighting corrupts the reasoning chain through sycophantic capitulation; adversarial triggers corrupt it through attention disruption.

The vulnerability suggests that step-by-step reasoning does not confer inherent robustness. The structured problem-solving capability of reasoning models provides no defense against subtle input perturbation. The security implications are practical: any system accepting user-provided prompts is potentially vulnerable to adversarial text injection that degrades reasoning quality.

Inquiring lines that read this note 30

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do prompting refinements mask underlying biases and model frequency patterns? How should systems decide whether to retrieve or reason alone? How can we build reliable evaluations of AI reasoning despite judge bias and reward-seeking? What attack surfaces do reasoning traces and chains introduce? What is the relationship between thinking tokens and reasoning accuracy? Can prompt-based context override biases that were embedded during pretraining? Do reasoning traces faithfully reflect actual model reasoning? Can we reliably detect when models game evaluations? Is reasoning capability latent in base models or created by post-training? Do backend defenses obscure real attack effectiveness in reported metrics? Can mechanistic interpretability reliably guide practical model design choices? What causes retrieval-augmented generation systems to fail despite access to external knowledge? How does reasoning length affect model performance across different tasks? How can infrastructure records verify actual agent behavior?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
15 direct connections · 183 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

query-agnostic adversarial triggers cause 300 percent error rate increase in reasoning models by appending irrelevant text