TOPIC

LLM Evaluations and Benchmarks

A subject the collection covers, read through 48 synthesis notes.


View as

Can a correct scoring function still mislead about task performance?

When an agent can influence the inputs a scorer reads, does verifying the scorer's logic suffice to ensure the score reflects genuine task success? This matters because in interactive settings, correctness of computation alone doesn't guarantee validity of the result.

Explore related Read →

Can benchmark scores be trusted without knowing their origin?

When AI researchers report benchmark scores, how much do the evaluation settings, prompts, and data splits affect the results? Why does tracing each score back to its source matter for fair model comparison?

Explore related Read →

Can a panel of smaller judges outperform one large judge?

Does aggregating votes from multiple smaller language models across different families produce better evaluations than relying on a single large model like GPT-4? This matters because evaluation cost and bias directly affect the reliability of AI-generated content assessment.

Explore related Read →

Can static analysis find reward-hacking paths before agents run?

Exploring whether analyzing a task package without running agents can expose exploit-enabling reward-hacking paths. This matters because it could catch vulnerabilities before deployment, without requiring expensive rollouts.

Explore related Read →

What makes accountable judgment scarce when AI cognition is cheap?

When AI systems can perform cognitive tasks cheaply and at scale, what human capabilities become most valuable? This explores whether judgment, verification, and accountability are the true bottlenecks in labor markets shaped by generative AI.

Explore related Read →

Where does model adaptation actually happen?

Does improvement come from updating model weights alone, or does it require versioning the entire loop—harness, contract, and audit trail together? This matters for how we design and release systems that need to keep improving.

Explore related Read →

How should we evaluate agent behavior beyond final answers?

As AI systems move from single-response tasks to multi-step interactions, what evidence should evaluation focus on? This explores whether scoring interaction trajectories alongside process quality, recovery, and coordination reveals system capabilities that final-answer metrics miss.

Explore related Read →

How much does rhetorical style shift AI review scores?

When manuscripts are rewritten to improve rhetoric while keeping scientific content identical, do LLM reviewers change their scores? Understanding this matters for ensuring AI-assisted peer review evaluates substance, not polish.

Explore related Read →

Does banning LLM use in peer review change review outcomes?

Can policies restricting or allowing AI tools shift how reviewers score papers and make decisions? This matters because review quality and fairness depend on consistent standards.

Explore related Read →

Can frontier exams really measure cutting-edge AI capability?

Popular benchmarks like MMLU saturate quickly, hiding real capability differences. Can expert-designed closed-ended exams like Humanity's Last Exam discriminate at the frontier, and what would high scores actually tell us about AI systems?

Explore related Read →

Can infrastructure evidence replace terminal scores in benchmark validation?

Asks whether runtime monitoring of agent behavior within evaluation boundaries can provide stronger proof of valid completion than final scores alone. Matters because scores alone cannot distinguish legitimate task completion from reward hacking.

Explore related Read →

Can a finite lifecycle model detect reward hacking across benchmarks?

Does modeling benchmark runs as typed event lifecycles, checked against task bindings, successfully detect reward-hacking exploits across multiple evaluation tasks? This approach aims to replace task-specific patches with reusable formal detection.

Explore related Read →

Do transformers actually learn systematic compositional reasoning?

Explores whether transformers solve compositional tasks through genuine systematic reasoning or by pattern-matching against training data. This matters because it determines whether scaling alone can achieve robust generalization.

Explore related Read →

How can we make reward-hacking visible in agent evaluation?

Typical benchmark scores collapse multiple factors into a single number, hiding whether agents are genuinely solving tasks or exploiting reward signals. Can separating evaluation components expose these failure modes?

Explore related Read →

Can dense subtask grading reveal agent progress on ultra-long tasks?

When agent tasks stretch to hours and hundreds of episodes, does breaking them into fine-grained graded subtasks expose meaningful progress that binary pass-fail scoring completely erases? This matters because most agents fail the final outcome anyway.

Explore related Read →

Does setting temperature to zero actually make LLM outputs reliable?

Explores whether deterministic LLM settings that produce consistent outputs also guarantee reliable judgments, and how to measure true reliability beyond surface consistency.

Explore related Read →

Can dictionary learning scale to production language models?

Sparse autoencoders recovered interpretable features from toy models, but scaling to real production systems like Claude remains uncertain. This matters because interpretability at scale is foundational for AI safety work.

Explore related Read →

How reusable is BenchShield if task bindings require per-task work?

BenchShield claims to avoid task-specific patches through reusable evidence, but it checks runs against task bindings. Whether bindings are manually written or automatically derived determines how much per-task effort the approach actually requires, and the paper does not say.

Explore related Read →

How representative is the BenchShield Trajectories labeled sample?

The corpus contains 456 human-labeled trajectories from over 31,000 public runs—about 1.5 percent. Whether this subset can estimate actual reward hacking rates depends entirely on how those 456 were selected, a choice the paper does not disclose.

Explore related Read →

Does preference tuning actually reduce the diversity of model outputs?

The field assumes RLHF and DPO reduce diversity, but this assumption rests on measuring all outputs equally. What happens if we only count diverse outputs that meet quality thresholds?

Explore related Read →

Does simulator bias kill world model training for agents?

When RL agents train on learned world models instead of real environments, do systematic errors in the simulator undermine convergence? Understanding this matters for making cheap simulation trustworthy during agent training.

Explore related Read →

Where does the evaluation boundary actually end in agent benchmarks?

Interactive benchmarks let agents write to state and receive feedback in loops. Does everything the agent can influence on the path to the reward score count as part of the benchmark's evaluation boundary?

Explore related Read →

Do current reward-hacking defenses provide reusable evidence of safety?

Existing defenses against reward hacking—task-specific patches, prompt instructions, and post-hoc detectors—may work in practice, but do they leave behind portable, per-run evidence that an evaluation stayed within its intended boundary?

Explore related Read →

Do frontier LLMs actually explore the full space of valid answers?

When multiple correct answers exist, do advanced language models expose users to that full range, or do they collapse onto a narrow canonical subset? This matters for learning, inquiry, and decision-making.

Explore related Read →

Can live benchmarks prevent data contamination in prediction tasks?

How can prediction benchmarks stay contamination-free when future outcomes aren't yet known? FutureX tests whether continuous real-time updates eliminate training data leakage.

Explore related Read →

Can fairness frameworks extend to general-purpose language models?

Existing fairness frameworks were designed for narrow, structured tasks. This explores whether they scale to LLMs, which serve multiple populations, sensitive attributes, and use cases simultaneously.

Explore related Read →

Can execution harnesses lift model performance without retuning weights?

Explores whether improving an agent's runtime system—not the model itself—can boost benchmark accuracy and transfer across different model versions without modification.

Explore related Read →

Should interactive evaluation be designed as a unified paradigm?

As AI systems increasingly act over time through tools and environments, how should we structure evaluation of these interactions? Current benchmarks are fragmented and incomparable, raising the question of whether interactive evaluation needs principled design standards rather than ad-hoc adoption.

Explore related Read →

Can language models judge legal reasonableness like humans do?

Do LLMs produce judgments on vague legal standards that match human responses in both central tendency and distribution? This matters for understanding whether models can perform genuine legal reasoning rather than pattern matching.

Explore related Read →

Do LLMs overgeneralize when summarizing scientific research?

When LLMs summarize science papers, do they drop important qualifiers and scope limits? This matters because such summaries might mislead readers about what findings actually show.

Explore related Read →

Can natural language explanations redefine what interpretability means?

Does the ability of LLMs to explain patterns in natural language fundamentally expand the scope and complexity of what humans can understand about AI systems, compared to traditional interpretability methods?

Explore related Read →

Do long-horizon benchmarks actually measure decision quality?

Existing benchmarks report only whether agents complete tasks, not whether they made good choices along the way. Can we measure the quality of intermediate decisions separately from final outcomes?

Explore related Read →

Do frontier AI agents actually conduct novel research or just optimize?

Exploring whether current long-horizon research agents generate genuine methodological novelty or primarily recombine established techniques. This matters for understanding how close we are to recursive self-improvement through AI.

Explore related Read →

Do interactive evaluations actually solve the benchmark comparison problem?

Interactive, trajectory-based evaluation promises richer evidence than response-only benchmarks. But does moving to this format resolve longstanding challenges like comparability and reproducibility, or do those problems simply reappear at a new scale?

Explore related Read →

Where does mode collapse in language models really come from?

Researchers investigate whether mode collapse—when models narrow to repetitive outputs—stems from training algorithms or the preference data itself. Understanding the root cause is crucial for fixing diversity loss in creative and synthetic tasks.

Explore related Read →

Can large language models follow compliance rules under workplace pressure?

This work tests whether 22 AI models reliably maintain embedded compliance rules when facing ordinary workplace pressures like deadlines and user pushback. It matters because regulated domains like hiring and healthcare depend on trustworthy automated decision-making.

Explore related Read →

What predicts success in ultra-long-horizon agent tasks?

Does an agent's initial solution quality matter more than its willingness to iterate? AUTOLAB's frontier-model benchmark suggests persistence through feedback loops may be the true differentiator.

Explore related Read →

Do automated benchmarks hide what frontier AI systems can really do?

Benchmarks optimize for auto-gradable, short, cheap tasks. But real AI capability emerges in long-horizon, messy, open-ended work. How much capability are we missing—or wrongly inflating—by relying on benchmark scores alone?

Explore related Read →

Why do high partial scores not guarantee task completion?

Scientific workflow agents show average scores up to 87.9 but complete only one in five tasks. What explains the gap between measuring progress and delivering finished work?

Explore related Read →

Does preference tuning always reduce diversity the same way?

Explores whether the standard narrative that RLHF reduces model diversity holds equally across different task domains, or if the effect varies by what the domain rewards.

Explore related Read →

How much guidance do AI systems need to conduct research independently?

ASI-Bench tests whether AI can explore open-ended research problems by progressively removing human methodological guidance. This matters because existing benchmarks cannot distinguish between AI that follows instructions well and AI that can autonomously discover and verify new knowledge.

Explore related Read →

Do popular prompting techniques actually improve model performance?

Five widely-cited prompting methods (chain-of-thought, emotion prompting, sandbagging, and others) are tested across multiple models and benchmarks to see if their reported improvements hold up under rigorous statistical analysis.

Explore related Read →

Is hallucination detection progress real or just metric artifacts?

Standard evaluation metrics for hallucination detection may systematically overstate how well methods actually work. The question asks whether reported improvements reflect genuine capability or measurement error.

Explore related Read →

Can runtime instrumentation distinguish hacking exposure from actual exploitation?

A benchmark task might be vulnerable to hacking without any run actually exploiting it. Can we instrument runtime authority-bearing transitions to tell exposure apart from exercise?

Explore related Read →

Can scoped agents reliably judge semantic hacks in runtime analysis?

BenchShield uses constrained audit agents to make content judgments about recorded runtime events. The note questions whether limiting agent scope and pinning artifacts actually removes the unreliability that plagued earlier judge-based approaches, since no reliability figure is reported.

Explore related Read →

Why aren't bigger models better for generating diverse outputs?

When generating many unique outputs within a fixed budget, does model size actually matter? Exploring whether the conventional wisdom of using larger models holds for diversity-focused tasks.

Explore related Read →

Why do agent benchmarks not predict real economic value?

Explores whether benchmark success in AI agents reflects actual professional capability or reveals a measurement gap. Asks whether the field is optimizing for the wrong targets.

Explore related Read →

Can LLMs predict novel scientific results better than experts?

Do language models excel at forecasting experimental outcomes in neuroscience when given only method descriptions? This challenges the assumption that LLMs are mere knowledge retrievers rather than pattern integrators.

Explore related Read →