Does setting temperature to zero actually make LLM outputs reliable?
Explores whether deterministic LLM settings that produce consistent outputs also guarantee reliable judgments, and how to measure true reliability beyond surface consistency.
"Can You Trust LLM Judgments?" (2024) introduces a rigorous framework for evaluating LLM-as-a-Judge reliability using McDonald's omega, revealing that the common practice of using fixed seeds and deterministic settings provides false confidence.
The core argument: even with deterministic settings, a single LLM output is one sample from the model's probability distribution. Setting temperature to zero and fixing the seed produces "fixed randomness" — the same output every time, but that output may still be a misleading draw from the distribution. Consistent replication does not guarantee reliability. A perfectly calibrated LLM that says it's 90% confident should be correct 9 out of 10 times — but even a perfectly calibrated LLM can be unreliable if its distribution has high variance.
The framework: prompt the judgment LLM 100 times, varying only the replication while holding all other factors constant. Apply McDonald's omega to assess internal consistency across these replications. This reveals whether the model's judgments are stable properties of the input or artifacts of the sampling process.
The distinction between reliability, confidence, and calibration is critical:
- Calibration: alignment between stated confidence and actual correctness
- Confidence: the model's self-assessed certainty
- Reliability: consistency of judgments across multiple draws
These three are intertwined but distinct. A model can be well-calibrated (confident when right) but unreliable (different answers on different draws). A model can be reliable (always gives the same answer) but poorly calibrated (that consistent answer is wrong).
This connects to Does model confidence predict robustness to prompt changes? — ProSA measures sensitivity to prompt variation, while this measures sensitivity to sampling variation. Both reveal that single evaluations are insufficient. The practical implication: any LLM-as-a-Judge deployment that relies on single-shot evaluation with deterministic settings is providing the illusion of precision without evidence of reliability.
Inquiring lines that read this note 169
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Why is hallucination an inevitable limitation of current language models?- What makes LLM outputs fabrication rather than hallucination or confabulation?
- What should we call errors in LLM outputs when hallucination does not apply?
- How much does ROUGE metric choice inflate hallucination detection claims?
- Does inevitable LLM hallucination make detection metric validity critical?
- What makes accountability and validity-orientation non-behavioral properties?
- Can systems lacking inner states express genuine truthfulness claims?
- Why does aggregate accuracy fail as a metric for rare harmful cases?
- Why do standard accuracy metrics ignore set-level consumption constraints?
- What makes the 45 percent accuracy saturation threshold universal?
- How should benchmarks balance verifiability against outcome resolution?
- Can a metric that rewards central tendency hide degenerate predictor failures?
- How can hidden test partitions detect constant predictions that generalize?
- What distinguishes minimal-pair asymmetry from standard accuracy evaluation?
- How should product specifications measure alignment without naming the dimension?
- What happens when alignment targets measure only the preferred dimension of entangled properties?
- Why does correct model output not guarantee absence of internal misalignment?
- What would it mean to assign explicit trust weights to synthetic data?
- How should ground truth labels be assigned to simulated user sessions?
- Can models detect statistical properties of their own generation in real time?
- How do users mistake synthetic LLM outputs for empirical observations?
- Can LLM judges reliably estimate when they lack sufficient persona information?
- What calibration corrections can reduce LLM judge bias in automated evaluation pipelines?
- What does McDonald's omega reveal about LLM judgment consistency?
- How do calibration and reliability differ in LLM judge evaluations?
- What other evaluation biases exist in LLM judge systems?
- Can crowdsourced voting and automated panels both credibly evaluate LLM outputs?
- Where should measurement systems sit to avoid recording bias?
- How sensitive are LLM bias measurements to analysis choices?
- How does unidimensionality in assessments affect measurement validity?
- What makes the Brier score mathematically better than log-likelihood here?
- Why is the Judging preference constant while other traits vary slightly?
- Why does sophisticated measurement not validate the underlying scientific inference?
- Can similar outputs from different systems prove they work the same way?
- Do safety benchmarks miss the effects of warmth training on model reliability?
- Can safety benchmarks detect reliability degradation from warmth training?
- Why do users systematically overrely on confident LLM outputs across languages?
- Can researchers prevent their expectations from shaping LLM outputs?
- Why does analytical depth demand trigger fabrication over transparent uncertainty?
- Can an LLM be well calibrated but still unreliable on single evaluations?
- Does exposure to more domain-specific examples reduce LLM overconfidence?
- How can we verify outputs from systems that generate without grounding?
- Which use cases can tolerate unverified LLM outputs without external verification?
- Why does regenerating LLM responses produce different but equally valid answers?
- Can users experience the LLM Fallacy even when AI outputs are completely accurate?
- What happens when we treat LLM outputs as sampled rather than stored?
- Can we systematically enumerate LLM failure modes from first principles?
- Do longer prediction horizons systematically degrade LLM forecasting accuracy?
- Can lightweight verification methods help experts trust LLM outputs?
- Why do LLM outputs need verification even when they look polished?
- How long does retrievability support error detection across repeated LLM use?
- What structural features force users to evaluate the epistemic status of outputs?
- What skills do users need to work effectively with stochastic outputs?
- What role does real-time accuracy feedback play in reducing user overreliance?
- Why does accumulated portfolio output not match accumulated worker capability?
- How does step-level confidence filtering compare to global confidence averaging?
- How do we assign confidence and polarity scores to belief edges?
- Does layer-wise prediction stabilization provide a stronger trace quality signal than confidence alone?
- Why does model confidence correlate with robustness to prompt variations?
- How reliable is the top-2 confidence gap as a stopping signal across tasks?
- Can uncertainty estimates based on model self-assessment reliably signal errors?
- Why do improvements in accuracy come at the cost of calibration?
- Does model confidence actually correlate with robustness against prompt variations?
- Can semantic entropy improve model calibration without external ground truth?
- Can proper scoring rules restore model calibration without sacrificing accuracy?
- What makes mathematically confident but incorrect answers resemble valid solution shapes?
- How does confidence in LLM outputs override users' ability to check accuracy?
- What makes uncertainty calibration harder than expanding knowledge?
- Can log-probability confidence be separated from decision-aligned signals?
- How do confident system outputs weaken user skepticism about their reliability?
- Why do fluent predictions fail to capture reliable internal models?
- What makes confident hallucination a distinct problem from poor calibration?
- Why is faithful calibration considered fundamentally metacognitive?
- Can distributional views explain when an LLM appears to change its mind?
- What property must remain constant to individuate an LLM across infrastructure changes?
- What distinguishes actual social disagreement from distributional uncertainty in LLM outputs?
- Can LLMs express uncertainty in ways that preserve epistemic honesty?
- What structural coherence exists in LLM preference systems and value hierarchies?
- How can we validate LLM-based drift measurements against human judgment?
- Can evaluation criteria be reliably encoded in labeled data without ground truth standards?
- When does the correlation between consistency and correctness break down?
- How should process quality and verification cost factor into evaluation judgment?
- Can log-likelihood loss combined with binary rewards achieve calibration?
- How does 93% reward reliability compare to other RL noise sources?
- Can utility control modify LLM values more effectively than output filtering?
- What consistency tests could distinguish constructed from genuine preferences?
- Why does preference measurement validity matter more than aggregation methods?
- Why does preference measurement validity matter before any aggregation?
- How do training data cutoffs produce false claims that stay consistent?
- What makes certain bond distributions more learnable than others?
- Can scaling predictions become reliable if improvements are continuous not sudden?
- How do surface statistical regularities enable correct outputs while degrading robustness?
- What makes output convergence across models inevitable given input-side homogenization?
- Why do models fail under distribution shift if accuracy metrics stay high?
- What consumption data would validate the limited-consumption model in production systems?
- Why do rare cases in medicine and science require models that preserve tail distributions?
- Can deterministic computation actually create new information in data?
- Can population-level distributions shift usefully even when individual prediction fails?
- Can defenders tighten the total-variation bound in practice with measured benign activation rates?
- What makes a bounded observer's ability to extract information different from apparent randomness?
- How does disembedding from social context collapse reliability despite factual accuracy?
- How much does omniscient evaluation overstate real-world simulation fidelity?
- How should monitoring intensity change based on task criticality?
- Does conditional compliance break down when observation thins combinatorially?
- Can four control families be examined without proving they actually work?
- Why do true and false LLM outputs use the same mechanism?
- Why do different LLMs converge on nearly identical outputs?
- Can measuring semantic entropy help us detect unreliable generations?
- Do high-disagreement items signal contested values or measurement noise?
- How do local soundness signals work across different problem domains?
- What makes some model capabilities reliable while others remain brittle?
- Why does externalized state beat parameter scaling for agent reliability?
- Why does MCP's portability come with determinism failures in production workflows?
- How does self-consistency compare to confidence as a proxy reward signal?
- Why does self-consistency fail as a proxy reward for correctness?
- What makes self-consistency a sufficient training target for the judge role?
- When does provable stability in latent dynamics fail to preserve fidelity?
- Can trustworthy scoring prevent persistent iteration from compounding errors?
- How does Goodhart's Law apply when safety measures become optimization targets?
- What makes uniform bounds the right choice for safety boundaries?
- Can LLMs recover true joint distributions from marginal census data?
- Can aggregate survey realism coexist with unreliable fine-grained effects?
- What role does human response variation play in LLM simulation accuracy?
- Can we detect superposition in LLM personality traits and stated preferences?
- Can explicit stress tests measure dispositional factors or only stimulus response?
- What distinguishes dynamic personality modeling from unreliable preference drift?
- What makes out-of-band monitoring better than in-band verification loops?
- What makes some analysis tasks stable enough for rigid generated interfaces?
- What makes a deployment paradigm credible for maintaining scientific integrity?
- Can imperfect uncertainty estimates still beat uniform oversight strategies?
- How do we measure marginal risk instead of speculating about misuse scenarios?
- Can a correct outcome hide a fundamentally unsound decision-making process?
- How can a single instrument measure errors across multiple system layers?
- How can distillation preserve uncertainty expression instead of optimizing it away?
- Can experimental outcomes be reliably distilled into reusable insights?
- How does distilling only inconsistent rollouts compare to distilling all generations?
- Can per-user adapters remain consistent without drifting or leaking?
- What trade-offs emerge between training objectives and model reliability?
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does model confidence predict robustness to prompt changes?
Explores whether a model's certainty about its answer determines how much it resists prompt rephrasing and semantic variation. This matters because it could explain why some tasks are harder to evaluate reliably.
prompt sensitivity and sampling sensitivity are complementary reliability concerns
-
Can LLM judges be fooled by fake credentials and formatting?
Explores whether language models evaluating text fall for authority signals and visual presentation unrelated to actual content quality, and whether these weaknesses can be exploited without deep model knowledge.
judge unreliability compounds with exploitable biases
-
Why do preference models favor surface features over substance?
Preference models show systematic bias toward length, structure, jargon, sycophancy, and vagueness—features humans actively dislike. Understanding this 40% divergence reveals whether it stems from training data artifacts or architectural constraints.
calibration failure at the preference model level adds to the reliability problem
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
- Can Large Reasoning Models Self-Train?
- LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails
- AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
- Using Large Language Models to Create AI Personas for Replication and Prediction of Media Effects: An Empirical Test of 133 Published Experimental Research Findings
- Can Machines Think Like Humans? A Behavioral Evaluation of LLM-Agents in Dictator Games
- DecepChain: Inducing Deceptive Reasoning in Large Language Models
- When Hindsight is Not 20/20: Testing Limits on Reflective Thinking in Large Language Models
Original note title
deterministic LLM settings create fixed randomness not reliability — a single output remains one draw from the model's probability distribution