Ordinary, Reasonable Chatbots: Do AI Models Track Human Legal Judgments?

Paper · arXiv 2609.06769 · Published September 6, 2026
LLM Evaluations and Benchmarks

As people increasingly rely on artificial intelligence (AI) for guidance in their own lives, scholars, lawyers, and even judges have begun to consider the role of AI in legal decision-making. As “silicon sampling”—the use of generative AI models in social science research—is now impacting academia, “silicon jurors” could make an appearance in courtrooms. This study joins an emerging line of research on generative AI models’ ability to simulate human legal judgments. In particular, we study how large language model (LLM)-powered chatbots respond to series of questions about legal reasonableness. When the law needs to judge the appropriateness of a behavior, it most often asks whether the behavior was “reasonable.” Yet despite the ubiquity of reasonableness judgments, they are the site of constant vexation for lawyers, judges, and lay people. Reasonableness seems inherently vague and unpredictable, since it relies on variable context and implicit conceptual schemas. Moreover, many scholars caution that reasonableness judgments may vary along demographic lines. We compare the answers of human participants to those of twenty-six LLMs across twenty-five different legally relevant reasonableness judgments.

Introduction. The popularity and fluency of generative AI chatbots like ChatGPT and Claude has led people to increasingly use these tools to answer a variety of different questions—about health, careers, relationships, and beyond. Given the perceived helpfulness and convenience of using chatbots in these domains, scholars, journalists, and even federal judges have contemplated the possibility that AI could aid in various legal decision-making tasks and even, potentially, replace juries. Just as “silicon sampling” is being pushed in the social sciences (Argyle et al. 2023; Dillion et al. 2023; Bisbee et al. 2024), “silicon jurors” may soon begin to impact the law (Grimm, Grossman, and Coglianese 2024). In this paper, we contribute to the emerging research on AI’s ability to simulate human legal judgments (Posner and Saran 2025). To do so, we assess how AI chatbots respond to a series of questions about legal “reasonableness.” When the law seeks to regulate the behavior of individuals or parties, the standard it most often reaches for is “reasonableness.”

Discussion / Conclusion. and Implications responses that are similar to human responses. Although the models often provided answers that were statistically different from humans, our total impression is one of overall coherence. When looking across the range of questions in our survey, the models’ responses did a strong job of approximating the human responses–even for an inherently vague legal standard. For no question is the mean or median response from the models wildly divergent from the human response. Simple visual analysis of the violin plots in Figure 1 indicates that both the central tendencies of the models and their overall distribution of responses tend to match the human reasonableness responses. We stress that the answers to legal reasonableness judgment questions are unlikely to exist in LLMs’ latent training data in the way that the answer to a legal question like “What is the minimum age for a U.S. Senator?” will be.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

What compositional reasoning failures limit large language models despite scale? Is language model reasoning authentic and what causes models to reason? Why is hallucination an inevitable limitation of current language models? What safeguards enable trustworthy AI-assisted scientific peer review at scale? How well do AI systems understand human social norms? What happens to knowledge when intelligence becomes tokenized like a commodity? Why does polished presentation create unearned authority in AI outputs? How does the generation-verification gap limit what we can measure about AI reasoning? What determines appropriate intervention timing and manner for AI agents? Why do some clarifying approaches produce understanding while others just satisfy? How do false presuppositions and sycophancy drive persistent false beliefs in models? How does evaluation scope and dimensionality affect what we measure?