SYNTHESIS NOTE
Topics›Prompts Prompting›this note

Does iterative prompt engineering undermine scientific validity?

When researchers repeatedly adjust prompts to get desired outputs, does this practice introduce hidden bias and produce unreplicable results? The question matters because LLM-based research is proliferating without clear methodological safeguards.

Synthesis note · 2026-03-28 · sourced from Prompts Prompting

"From Prompt Engineering to Prompt Science" (2024) argues that using LLMs for scientific research through iterative prompt revision is methodologically dangerous. The standard practice — a researcher iteratively tweaks prompts until the LLM produces desired outputs — violates core scientific principles.

Three specific problems:

Individual bias and subjectivity. When a single researcher revises prompts ad hoc, personal biases shape the prompt trajectory. The researcher's expectations about what constitutes a "good" output steer the revision process, potentially embedding those expectations into the prompt without explicit awareness or documentation.

Vague or shifting criteria. Without pre-specified evaluation criteria, the definition of a "desirable outcome" drifts during prompt revision. Worse, researchers may unconsciously bend criteria to match what the LLM can produce, rather than holding the LLM to task-appropriate standards. This is a form of overfitting hypotheses to data.

Self-fulfilling prophecy. The opacity of LLMs makes feedback loops especially dangerous. A prompt revised to produce output that "looks right" may be finding a local optimum in the model's generation space that happens to align with the researcher's expectations — not because the underlying task is being solved correctly, but because the prompt has been tuned to produce the expected surface pattern.

The proposed alternative adapts qualitative coding methodology: (1) at least two qualified researchers, (2) pre-specified, community-validated evaluation criteria before any prompt revision, (3) explicit discussion of individual differences and biases, (4) iterative development until inter-coder reliability (ICR) is achieved, (5) independent validation on unseen data by different researchers.

The practical validation: applying this pipeline to user intent taxonomy construction with GPT-4 achieved high ICR between human annotators and between humans and LLM. However, the authors note a fundamental limitation: "even with enough documentation of how the final prompt was generated, one could not ensure that the same process will yield same quality output in a different situation — be it with a different LLM or the same LLM in a different time."

Foundation Priors formalization. The Foundation Priors paper (2024) provides a formal statistical framework for these dangers. It models prompt engineering as an iterative alignment process where users minimize divergence between synthetic output and their anticipated data distribution. The self-fulfilling prophecy is, formally, epistemic circularity: the user refines until the output matches their priors, then treats the match as evidence their priors are correct. The paper introduces a trust parameter λ that governs how much weight synthetic data should receive in inference — making explicit what ad-hoc prompt engineering leaves implicit (full trust, λ=1). Since Should we treat LLM outputs as real empirical data?, the Foundation Priors framework upgrades the methodological critique here into a formal epistemic one: the problem is not just unreliable method but miscategorized output.

This connects directly to the custodial shift. Since How does LLM-mediated search change what expertise requires?, the expert's new responsibility includes not just prompt skill but prompt rigor. Ad-hoc prompting is the custodial equivalent of running uncontrolled experiments — it may produce useful results, but it cannot produce trustworthy ones. The shift from engineering to science mirrors the broader shift from producing knowledge to validating AI-generated knowledge.

Inquiring lines that read this note 27

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do prompt design choices influence model reasoning and performance? How do prompting refinements mask underlying biases and model frequency patterns? Why don't LLMs reliably translate capability into accurate outputs? Is language model reasoning authentic and what causes models to reason? Why can't prompting alone inject genuinely new knowledge into models? Can brute-force automated research substitute for iterative depth and human research intuition? How does evaluation scope and dimensionality affect what we measure? What safeguards enable trustworthy AI-assisted scientific peer review at scale? What do systematic disagreements between annotators reveal about ground truth? How do LLM judges' systematic biases affect alignment and evaluation outcomes?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
17 direct connections · 130 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

ad-hoc prompt engineering violates scientific method — producing unreliable biased and unreplicable research outcomes that risk self-fulfilling prophecy