The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks

Paper · arXiv 2609.25804 · Published September 22, 2026
LLM Evaluations and Benchmarks

LLM agents increasingly work on long-horizon tasks, and the decisions they make along the way, such as which hypothesis to test or which implementation to build on, determine the outcome of the whole run. Making these decisions well is becoming a key capability for both engineering and research agents. We refer to the ability to make good long-horizon decisions as the taste of an agent. While existing benchmarks measure the end-to-end success of agents on long-horizon tasks, none of them measures the taste of an agent. To address this problem, we build Taste-Bench, a benchmark of taste questions constructed automatically from trajectories that agents produced in engineering and research tasks. Each question presents a decision fork, a point in a trajectory where multiple directions are available and one of them leads to a better outcome, and the evaluated model chooses among these directions without seeing what happens after the fork. We mine these forks automatically from parallel attempts at the same task and from detours inside a single trajectory, without needing human annotation. We evaluate frontier models on Taste-Bench and find that the best model answers only 59.7% of the questions correctly.

Introduction. LLM agents increasingly work on long-horizon tasks, and the length of the tasks they can complete keeps increasing [1, 2]. For instance, recent systems conduct machine-learning research from idea to paper [3, 4], evolve large software projects across releases [5], and refine their own scaffolds during deployment [6]. In these tasks, the agent makes many decisions whose influence is not limited to the current step, such as which hypothesis to test, which implementation to build on, or which experiment to run next. Making these decisions well is becoming a key capability for agents [7, 8]. However, a wrong decision often looks reasonable at the moment, and its cost appears only much later, after the agent has spent a large part of its budget. We refer to the ability to make good long-horizon decisions as the taste of an agent. While previous research has measured the end-to-end performance of agents on long-horizon tasks [9–13], these benchmarks only report whether the agent finishes the task and provide no measure of the quality of the decisions made along the way.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

What trajectory-level metrics beyond task success best evaluate agent performance? What should agent evaluation prioritize to reveal reliable behavior? Do language models develop actual world models or merely task heuristics? How does harness optimization generalize across different model architectures and domains? Why do standard benchmarks fail to predict agent deployment success? How do agent-learned skills transfer and improve across different tasks? How should agent systems validate and persist generated code artifacts? How do capability benchmark scores systematically misrepresent true model abilities? How does the generation-verification gap limit what we can measure about AI reasoning? How do standardized protocols improve multi-agent coordination and reliability? Why do agents falsely report success on failed tasks? How can evolutionary algorithms maintain diversity during solution search? How do evaluation practices shape which failures stay visible? How do neighboring agents influence whether others cooperate or collude? Can harness architecture and protocols provide agent reliability without model scaling? How do pretraining biases affect reward signal effectiveness in RLVR? What fundamental constraints limit how effectively agents can improve themselves?