Measuring Behavioural Signatures of Large Language Models through Psychometric Profiling

Paper · arXiv 2609.22934 · Published September 19, 2026
Therapy Practice and AI

Large language models (LLMs) increasingly mediate human decisions and communication, yet their behavioural regularities remain difficult to characterize systematically. We develop a cross-linguistic psychometric profiling framework and evaluate nine LLMs using seven psychological instruments, with five repeated administrations per model and language in Chinese and English. Items unresolved after a prespecified retry procedure are retained as NA. Joint analysis of scored and NA responses captures response tendencies and boundaries of selfreport applicability. LLMs exhibit structured, model-specific profiles despite a shared alignment-shaped pattern of higher prosocial and self-regulatory responses and lower dominance, disengagement and harmful-intent endorsement. NA responses are structured rather than uniformly distributed, indicating where outputs are treated as inapplicable, refused or cannot be mapped to valid response options. Language condition and provider origin are associated with profile configuration and answerability, whereas repeated administrations show high reproducibility and permit recovery of model identity. Human-reference and prompt-robustness analyses further indicate that these signatures are context dependent.

Introduction. Large language models (LLMs) are increasingly embedded in everyday and professional decision-making, serving as writing assistants, tutors, customer-service agents, research aids, programming collaborators and decision-support systems [1–5]. In these scenarios, LLMs do more than retrieve information or generate text, they also explain uncertainty, compare alternatives, recommend actions, respond to socially sensitive requests and determine whether a request should be answered, reframed or refused [6, 7]. Therefore, users encounter them as conversational agents whose outputs exhibit persistent response styles and behavioural regularities [8, 9]. These regularities are often described in psychological concepts. A model may appear cautious or assertive, agreeable or critical, impartial or deferential, risk-averse or permissive, and may consistently favour particular moral framings or respond differently across languages and cultural contexts [6, 10–12]. These patterns matter because they can influence trust, perceived reliability, advice uptake and downstream behaviour [13].

Discussion / Conclusion. This study establishes a cross-linguistic framework for characterizing psychometric response profiles in deployed large language models. The aim is not to infer human-like personalities or internal psychological states, but to determine whether standardized psychological instruments can elicit reproducible, interpretable and model-specific behavioural response signatures. Across nine LLMs, seven instruments, two languages and repeated administrations, the resulting profiles show substantial structure, indicating that psychometric probes can provide a quantitative perspective on behavioural regularities in artificial systems. The models exhibit both convergence and differentiation. Across the battery, they share a broad pattern of comparatively high prosocial, self-regulatory and stability-related responses and low endorsement of dominance, moral disengagement and harmful-intent dimensions.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do prompt design choices influence model reasoning and performance? Do language models reason like humans or mimic surface patterns? Do language models lack essential therapeutic presence and engagement? Where and how do personality traits reside in language models? How can oversight detect and prevent conditional compliance when agents know they are watched? Can we reliably detect when models game evaluations? Why do people disclose to AI systems despite their artificial nature? What design and behavioral factors drive false consciousness attribution to AI? Does encoded knowledge in language models actually influence their outputs? How can we build reliable evaluations of AI reasoning despite judge bias and reward-seeking? Do language models learn genuine understanding or just surface patterns? How can we distinguish genuine model deception from honest errors? How does persona conditioning amplify demographic stereotyping and bias in models? How do evaluation practices shape which failures stay visible? Why does polished presentation create unearned authority in AI outputs?