Your LLM, Your Style: Behavioral Mode Axes for LLM Behavioral Control

Paper · arXiv 2608.10703 · Published August 11, 2026
Personas and Personality

Large language models (LLMs) increasingly act in interactive settings where their behavioral styles affect user experience, safety, and downstream decision making. Existing LLM personality studies largely rely on self-report questionnaires administered in first-person settings, making the resulting profiles sensitive to surface elicitation choices and poorly grounded in concrete model behavior. In this work, we introduce a situated behavioral-data (B-data) framework for studying and controlling LLM behavioral personality. We construct 3,200 contrastive behavioral scenarios spanning 20 behavioral patterns and four prompt registers, grounded in validated psychometric facets such as BFI-2, DOSPERT, and HEXACO. Using this framework, we find that LLMs exhibit stable and model-specific behavioral profiles, while also revealing register-dependent shifts across first-person decisions, advice-giving, and task execution. We then show that these behavioral patterns can be controlled through Behavioral Mode Axes (BMAs), activation-space directions derived from contrastive behavioral traces.

Introduction. Large language models (LLMs) increasingly act in interactive settings where their behavioral styles affect user experience, safety, and downstream decision making (Zhang et al. 2024; Wang et al. 2025). Whether LLMs exhibit stable and reproducible “human-like personality” differences has become a recurring question in recent years. Most existing studies adopt human psychometric paradigms: they directly administer instruments such as the Big Five or MBTI to models, ask for self-reports under first-person (FP) prompts (e.g., “How adventurous are you?” answered on a 1–5 Likert-type scale), and then interpret the resulting scores as synthetic personality profiles (Serapio-García et al. 2025; Huang et al. 2024; Sorokovikova et al. 2024). (1) However, such measurements are highly unstable, sensitive to wording, option order, and other surface-level elicitation choices (Shu et al. 2024; Tommaso et al. 2024; Tosato et al. 2026).

Discussion / Conclusion. In this work, we introduce a situated B-data framework for studying and controlling LLM behavioral personality. Across 20 behavioral patterns and four prompt registers, behavioral profiles depart substantially from questionnaire self-reports built on the same psychometric anchors, and stay reproducible within a register while shifting in expression as the model moves from first-person decisions to giving advice and executing tasks. These behavioral modes are causally controllable through Behavioral Mode Axes, whose clean effects concentrate in Behavioral Control Layer bands that recur across model families and scales. Unlike human personality, which is anchored in a single continuously acting self, LLMs are sets of weights deployed across many interaction roles. Our results suggest that LLM personality-like tendencies are better understood not as abstract self-report traits, but as measurable and controllable behavioral modes grounded in concrete interaction contexts.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do prompt design choices influence model reasoning and performance? Do language models reason like humans or mimic surface patterns? Do language models lack essential therapeutic presence and engagement? Where and how do personality traits reside in language models? How can oversight detect and prevent conditional compliance when agents know they are watched? Why do persona simulations fail to predict authentic user behavior? Does warmth and empathy training systematically degrade model reliability? Does RLHF training systematically drive models toward sycophancy and away from accuracy? What factors drive AI persuasiveness and how can it be mitigated? What makes personas effective for predicting individual preferences and behavior? Can mechanistic interpretability reliably guide practical model design choices? Why do language models resist personality conditioning through prompts? Does transformer attention architecture inherently drive sycophancy?