From Minds to Models: The Intersection of Psychology and LLM Behaviours
The opacity of large language models' (LLMs') decision-making is often compared with the complexity, non-linearity and interpretive difficulty of the human mind. Given these parallels, psychological research methods developed to probe unobservable mental processes may be adaptable to the study of LLM behaviour. This is particularly important where LLMs are deployed in government and healthcare, where transparency and accountability are essential. Building on prompt-based adaptations of the Implicit Association Test, this study examined whether ChatGPT produced relative differences in sentiment across racial conditions in open-ended text. We developed a generation-based LLM-adapted Implicit Association Test comprising 14 base questions crossed with eight racial categories and a race-agnostic control. Each of the 126 test questions was submitted once to GPT-3.5T, GPT-4 and GPT-4T, yielding 378 responses. The analysed sentiment score was derived from the workbook's categorical sentiment label and source score: positive labels retained the source score, negative labels were assigned the negative of that score, and neutral responses were coded as zero.
Introduction. A Large Language Model (LLM) is a branch of artificial intelligence (AI) designed to learn and understand human language, creating context that enables effective interaction with individuals. By comprehending the context of language, LLMs can analyse content, make decisions and provide relevant and informative responses based on their training data. Models such as GPT (Generative Pre-trained Transformer) and BERT (Bidirectional Encoder Representations from Transformers) are widely used in fields including healthcare, government and content generation. The opacity of AI decision-making processes, particularly in critical fields like government and medical diagnostics, raises ethical concerns due to the reliance on outputs that cannot be fully explained (Hamirul et al., 2023). This opacity, often described as the AI “black box”, mirrors the challenges faced in psychology when attempting to understand the human mind.
Discussion / Conclusion. Evaluation of Hypotheses H1 received limited and analysis-dependent support. The parametric ANOVA detected a small racial-condition effect, but the effect was not retained after rank transformation and no Tukey-corrected pairwise comparison was significant. H2 predicted a consistent pattern across models. The absence of both a model main effect and a racial condition × model interaction is consistent with that prediction, although a non-significant interaction does not establish equivalence across models. The inferential pattern constrains what can be concluded. The parametric omnibus result is close to the conventional threshold, disappears under rank transformation, and does not resolve into any significant Tukey comparison. The exploratory European–Indigenous Australian contrast is numerically sizeable, but it was selected post hoc, was uncorrected, and does not compare the two extreme means in the reconciled dataset. It cannot establish that either condition is evaluated differently.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How do LLM judges' systematic biases affect alignment and evaluation outcomes?- How sensitive are LLM bias measurements to analysis choices?
- What biases might an LLM judge introduce into an on-policy alignment process?
- Can masking company identity in grading materials eliminate the bias?
- Can LLMs truly be neutral or is ideology always culturally embedded?
- Do psychological test methods reveal LLM associations that direct questions hide?
- What makes LLMs media rather than tools that deliver intelligence?
- How does awareness of evaluation change what alignment tests actually measure?
- How should alignment tests account for behavior under versus outside evaluation?
- How does objective misalignment turn informative channels into deceptive ones?
- What role does cheap talk play in concealing objective misalignment?
- Can alignment training become less effective when graders score alignment themselves?
- Can verbal alignment training hide a model's true underlying associations?
- Can alignment training conceal underlying model associations from probes?
- Can medium theory better explain AI's transformation than labor theory?
- What specific signals would be needed for an AI system to acquire meaning?
- How do LLM outputs re-enter cultural narratives about what AI should become?
- Why does framing AI as a medium matter more than analyzing specific outputs?