How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review
As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when reported scientific content is preserved and how these effects vary across evaluation conditions. We construct a controlled corpus of 4,200 full-paper manuscripts derived from 120 anonymized ICLR 2026 submissions. Two LLM rewriters transform six rhetorical dimensions in opposing directions, and five LLM reviewers evaluate the resulting manuscripts under standard and strict protocols. We also test joint, recursive, and reviewer-guided rewriting. Our results show that rhetorical sensitivity is structured rather than uniform. Evidence framing and novelty stance produce the largest positive-negative contrasts in overall assessment, with scope framing forming a weaker second tier; the remaining dimensions have smaller or less stable effects. This hierarchy persists across human-assessed quality levels, but score movement depends strongly on the AI reviewer’s original score: lower scores tend to rise, higher scores tend to fall, and directional contrasts are clearest in the middle ranges. More elaborate workflows do not reliably yield larger gains.
Introduction. Large language models (LLMs) are increasingly involved on both sides of scientific evaluation: they can revise how authors present their work and serve as scalable evaluators that accelerate review and reduce reviewer burden (Wang et al., 2020; Liang et al., 2024; Thakkar et al., 2026; Chen et al., 2026; Kaneko, 2026; Baumann et al., 2026). As these uses meet in the same evaluation pipeline, a central concern is how rhetorical presentation can reward-hack AI reviewers by changing their judgments without corresponding improvements in the underlying science. Presentation is especially relevant because the same scientific work can be communicated through different rhetorical choices: supported claims may be framed more assertively or cautiously, reported evidence may receive greater or lesser emphasis, and technical content may be expressed with different levels of complexity.
Discussion / Conclusion. This paper shows that AI scientific review is systematically sensitive to rhetorical presentation even when reported scientific content is preserved. This sensitivity is not uniform. It depends on the rhetorical dimension, the rewriting process, the rewriter and reviewer models, and the review protocol. The resulting variation cannot be reduced to a single average effect or treated as a stable property of one model configuration. These findings suggest that AI-assisted review should be evaluated for rhetorical robustness across multiple models and conditions, rather than judged solely by aggregate agreement or average scoring behavior. This study has several limitations. First, the benchmark is restricted to ICLR 2026 submissions with recoverable full-paper source and public review metadata, so the results may not generalize to other venues or scientific fields. Second, full-paper rewriting and multi-model evaluation are computationally expensive.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
What safeguards enable trustworthy AI-assisted scientific peer review at scale?- How do LLM reviewer scores respond when rewriting is applied recursively or jointly?
- Why do evidence framing choices move AI review scores more than other rhetorical changes?
- What prevents scholarly infrastructure from filtering out ghost-authored records automatically?
- How can automated review scale with the flood of AI-generated papers?
- What accountability structures should replace detection when AI automation increases in peer review?
- What collaboration model between humans and AI best serves peer review?
- How much has peer review workload grown at major conferences?
- Can automated AI systems assess novelty as well as human reviewers?
- Why should AI research prompts be subject to peer review before use?
- At what collaboration level should AI reviewers make final acceptance decisions?
- What discovery accuracy would satisfy the false-alert workload reviewers can tolerate?
- How does machine feedback enable discovery at test time?
- Can verification and accountability sustain meaningful human work at scale?
- Why does verification of AI work consistently lag behind AI generation?
- Can publishing failure branches change incentives to expose messy research processes?
- Why is verification harder than generation across the research lifecycle?
- How do closed-loop automated venues differ from human-in-the-loop review taxonomies?
- How does rubber-stamping differ from loss of scrutiny capacity in review processes?