Towards Scalable Measurement of Durable Skills

Paper · arXiv 2609.15864 · Published September 14, 2026
Argumentation and Persuasion

Durable skills, such as collaboration, creativity and critical thinking, are instrumental to success in the modern workforce. Yet, measuring these skills remains a persistent challenge. Moreover, because what is not measured is often not taught, these skills are often overlooked in mainstream educational curricula. Designing effective assessments for these skills necessitates balancing two often-conflicting requirements: ecological validity and psychometric rigor. On the one hand, the assessment environment should emulate natural interaction between humans, which is how these skills will be performed in the real world. On the other hand, it should be scalable, controllable and reproducible. Here we argue that Large Language Models (LLMs) can be used to better capture both of these aims. Concretely, we develop an AI-based framework where the subject converses with AI teammates in a way that resembles human-human interaction for authenticity, while also offering the psychometric control required for informative and robust assessment. Importantly, the AI participants not only act as teammates but also, in an “Executive LLM” setup, steer the conversation towards eliciting a high density of observable evidence for skill proficiency.

Introduction. Success in the modern workplace requires not only technical knowledge and procedural fluency, but additionally a host of future-ready human skills, including collaboration, communication, critical thinking, and creativity [1, 2]. Often also referred to as 21st century skills [3–7], these constructs remain notoriously hard to measure [8–10]. Previous efforts to resolve these measurement challenges and align educational goals, assessment, and instruction have focused on technological innovation, for example through automated scoring of answers [11–13] and the collection and analysis of detailed process data [14, 15]. Progress in large language models (LLMs) has opened up a new technological frontier, which we pursue here. Our key idea is that LLMs can bridge the gap between unstructured student collaboration, which more closely emulates classroom practice, and standardized assessment, which, while artificial, attempts to isolate the behaviors needed for valid inference.

Discussion / Conclusion. We have argued that large language models have the potential to transform the assessment of complex durable skills. In the case of group work, they help overcome a number of fundamental psychometric challenges regarding the reliability, comparability, scalability, and ecological validity of such assessments. Until now, these challenges have been mostly addressed by highly scripted interactions with AI teammates (e.g., PISA 2015) or highly structured human-human interactions (ATC21S). The novel introduction of the Executive LLM allows standardization of the collaboration experience, without overly scripting the interaction itself. Our results for collaboration also demonstrate that LLMs can be used to score conversations according to a rubric, and their agreement with human raters is similar to inter-rater agreement between humans. These results join a growing body of work showing that LLMs can be used to considerably scale the assessment process [44–46].

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

What prevents conversational agents from taking initiative in dialogue? How do training data properties determine the emergence of internal misalignment? How can AI chatbots provide therapeutic benefit without causing harm? Do reasoning benchmarks predict model performance in long-horizon workflows? Why don't LLMs reliably translate capability into accurate outputs? Can multi-agent systems avoid converging on false agreement without deliberation? Do language models reason like humans or mimic surface patterns? Do language models respond to social pressure and face-saving like humans? How do multi-agent LLM systems fail distinctly compared to single agents? How does dialogue structure affect linguistic grounding and shared meaning? Why doesn't reasoning volume improve theory of mind performance? How do neighboring agents influence whether others cooperate or collude? Can language models build genuine grounding through interaction? Do language models reason through causal mechanisms or semantic associations?