Natural Language Processing Psychometrics

Paper · arXiv 2608.07316 · Published August 7, 2026
Therapy Practice and AI

Natural Language Processing (NLP) models predicting mental health outcomes rarely specify what they measure: contextual knowledge, emotional content, or syntactic structure. NLP Psychometrics treats psychological prediction from text as a psychometric problem, linking scores to interpretable linguistic evidence and testing beyond the training text format. Nine LLMs, conditioned on controlled personas (cognitive digital shadows), completed psychometric questionnaires with textual explanations per item. We extracted emotional profiles and syntactic-semantic structure via textual forma mentis networks, combined with personality and sociodemographic variables in ablated random forest (RF) regressors, using SHAP to identify which features drove performance and in which direction. Full RF models explained up to 70.8% of variance in life satisfaction (SWLS), 55.7% in depression (PHQ-9), and, for DASS-21, 68.5% depression, 76.0% anxiety, 72.4% stress. Sociodemographics alone explained no meaningful variance in depression, anxiety, or stress, but did so for life satisfaction, where emotion features and income were the strongest predictors; neuroticism and network topology instead dominated depression and anxiety, reversing direction between them.

Introduction. Measuring psychological phenomena depends on inner experience leaving quantifiable traces behind [1–3]. Some of those are explicit, as in questionnaire ratings: Individuals consciously rate experiences and leave correlated numerical sequences, or item scores, as traces of their inner world [1, 3]. Other traces are linguistic [4–6], distributed across the words people choose, the emotions they express, and the concepts they connect, either explicitly or implicitly. In this view, language is not only a vehicle for communication [7, 8] but a key trace for psychological measurement, whose structured knowledge can open the way to psychological measurements [9, 10]. This premise is deeply connected to research on the mental lexicon [8, 11, 12]. In the modelling metaphor of the mental lexicon, language is reflected within cognition as a complex system of interconnected concepts or words [13]. The latter are not independent labels attached to experience, but structured cognitive representations embedded in networks of meaning [8].

Discussion / Conclusion. From our pioneering work with NLP Psychometrics, three findings stand out. First, language carried most of the recoverable psychometric signal: emotion and network features formed a predictive core of machine learning features across all tested scales (SWLS, PHQ-9 and DASS-21), whereas sociodemographics alone rarely explained meaningful variance. Second, the markers identified by SHAP scores were interpretable and construct-specific, ranging from family income and affect for life satisfaction to neuroticism and discourse topology for depression. Third, the machine learning "feature to psychometric score" mapping transferred, with reduced yet significant accuracy, to both out-of-genre LLM-generated diaries and to human clinically labelled data [75, 77], i.e., speech transcripts of clinically depressed patients and controls.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Can AI systems distinguish genuine empathy from simulated emotion? Do language models lack essential therapeutic presence and engagement? Can compression size predict model complexity better than parameter count alone? Where and how do personality traits reside in language models? How does persona conditioning amplify demographic stereotyping and bias in models? Do writers recognize when AI writing assistance alters their expressed stance? How does synthetic data quality and diversity affect downstream model capabilities? What explains language models' asymmetric difficulty with implicit versus explicit linguistic relations? Why do persona simulations fail to predict authentic user behavior? Why do some clarifying approaches produce understanding while others just satisfy? Do language models reason through causal mechanisms or semantic associations? Does model confidence reliably signal actual accuracy in practice?