SYNTHESIS NOTE
Topics›this note

Can we defend RAG systems from corpus poisoning without retraining?

Explores whether retrieval-time defenses can catch and block poisoned documents before they reach the generator, without expensive retraining cycles. Matters because corpus updates outpace model retraining in production RAG systems.

Synthesis note · 2026-05-03

RAG poisoning attacks insert malicious documents into the retrieval corpus so they get pulled in for matching queries and steer generation toward attacker-preferred outputs. Existing defenses typically require retraining the retriever or the generator, which is expensive and slow to deploy. RAGPart and RAGMask propose two lightweight defenses that operate at retrieval time without modifying the generation model.

RAGPart exploits a structural property of dense retrievers: they learn discriminative patterns from how the training data is partitioned, which means malicious documents inserted into one partition have predictably limited influence on retrieval from queries that match a different partition. By configuring partitions deliberately, the system bounds how far any single poisoned document can propagate. RAGMask takes a different angle: it masks tokens in candidate documents and watches for abnormal similarity shifts. Genuine documents are robust to token masking — their similarity scores degrade smoothly — while poisoned documents that rely on specific trigger tokens show sudden similarity collapse, which serves as a detection signal.

The architectural significance is that defense need not be coupled to training. Both methods sit at the retrieval layer and treat the generator as an untrusted black box that must be protected from upstream corruption. This separation matters operationally because retrieval corpora update faster than retrievers can be retrained, so defenses that require retraining are always behind the threat. The threat surface is real and severe — How vulnerable is GraphRAG to tiny text manipulations? shows even minimal corpus modifications can devastate accuracy in graph-structured RAG.

Inquiring lines that read this note 56

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do training data properties determine the emergence of internal misalignment? What causes retrieval-augmented generation systems to fail despite access to external knowledge? What attack surfaces do reasoning traces and chains introduce? When do semantic similarity approaches miss structural retrieval failures? How can we prevent synthetic data from contaminating statistical inference and corpora? How do surface patterns enable correct outputs but reduce robustness? Can inoculation prompting prevent emergent misalignment after reward hacking? Do knowledge graphs offer advantages over embeddings for multi-hop retrieval? Can prompt-based context override biases that were embedded during pretraining? Why does adding new knowledge through fine-tuning degrade existing capabilities? Can single-point security defenses protect multi-agent systems from multi-step attacks? Do backend defenses obscure real attack effectiveness in reported metrics? How should agents manage memory granularity to improve long-term performance? What fundamental constraints limit how effectively agents can improve themselves? Can we reliably detect when models game evaluations? How should retrieval systems handle complex multi-step reasoning? What safeguards enable trustworthy AI-assisted scientific peer review at scale? How can we build reliable evaluations of AI reasoning despite judge bias and reward-seeking? How effective are honeytokens and decoys against different security threats? Why do locally safe actions create system-level safety gaps? Can causal models help detect and locate hidden sandbagging in AI? How can we detect and prevent harm propagation through multi-agent delegation workflows? How vulnerable are token issuance and authorization policies to coordinated attacks? How can infrastructure records verify actual agent behavior? Do reasoning benchmarks predict model performance in long-horizon workflows?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
15 direct connections · 149 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

RAG corpus poisoning has lightweight defenses without retraining — partition-aware retrieval and token-masking similarity shifts catch attacks the generator never sees