Your LLM, Your Style: Behavioral Mode Axes for LLM Behavioral Control
Large language models (LLMs) increasingly act in interactive settings where their behavioral styles affect user experience, safety, and downstream decision making. Existing LLM personality studies largely rely on self-report questionnaires administered in first-person settings, making the resulting profiles sensitive to surface elicitation choices and poorly grounded in concrete model behavior. In this work, we introduce a situated behavioral-data (B-data) framework for studying and controlling LLM behavioral personality. We construct 3,200 contrastive behavioral scenarios spanning 20 behavioral patterns and four prompt registers, grounded in validated psychometric facets such as BFI-2, DOSPERT, and HEXACO. Using this framework, we find that LLMs exhibit stable and model-specific behavioral profiles, while also revealing register-dependent shifts across first-person decisions, advice-giving, and task execution. We then show that these behavioral patterns can be controlled through Behavioral Mode Axes (BMAs), activation-space directions derived from contrastive behavioral traces.
Introduction. Large language models (LLMs) increasingly act in interactive settings where their behavioral styles affect user experience, safety, and downstream decision making (Zhang et al. 2024; Wang et al. 2025). Whether LLMs exhibit stable and reproducible “human-like personality” differences has become a recurring question in recent years. Most existing studies adopt human psychometric paradigms: they directly administer instruments such as the Big Five or MBTI to models, ask for self-reports under first-person (FP) prompts (e.g., “How adventurous are you?” answered on a 1–5 Likert-type scale), and then interpret the resulting scores as synthetic personality profiles (Serapio-García et al. 2025; Huang et al. 2024; Sorokovikova et al. 2024). (1) However, such measurements are highly unstable, sensitive to wording, option order, and other surface-level elicitation choices (Shu et al. 2024; Tommaso et al. 2024; Tosato et al. 2026).
Discussion / Conclusion. In this work, we introduce a situated B-data framework for studying and controlling LLM behavioral personality. Across 20 behavioral patterns and four prompt registers, behavioral profiles depart substantially from questionnaire self-reports built on the same psychometric anchors, and stay reproducible within a register while shifting in expression as the model moves from first-person decisions to giving advice and executing tasks. These behavioral modes are causally controllable through Behavioral Mode Axes, whose clean effects concentrate in Behavioral Control Layer bands that recur across model families and scales. Unlike human personality, which is anchored in a single continuously acting self, LLMs are sets of weights deployed across many interaction roles. Our results suggest that LLM personality-like tendencies are better understood not as abstract self-report traits, but as measurable and controllable behavioral modes grounded in concrete interaction contexts.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How do prompt design choices influence model reasoning and performance? Do language models reason like humans or mimic surface patterns?- Why do questionnaire-based personality scores fail to predict actual LLM behavioral choices?
- Do psychological test methods reveal LLM associations that direct questions hide?
- How can we validate LLM-based drift measurements against human judgment?
- How does personality priming change LLM strategic decision making?
- Does the same linguistic signal work across patient speech, LLM text, and diary entries?
- Can personality control improve training outcomes for crisis workers and therapists?
- How does language condition affect model psychological profile consistency?
- What distinguishes dynamic personality modeling from unreliable preference drift?
- Do open language models default to a single shared personality type?
- How do LLMs identify which personality items matter most for trait inference?
- Why can data filtering fail to remove transmitted behavioral traits?
- Can continuous persona vectors in activation space monitor personality shifts?
- Do personality traits occupy specific mechanistic locations in pretrained models?