Interpreting and Steering LLM Agents for Social Simulations

Paper · arXiv 2609.16436 · Published September 14, 2026
Dialog Topics and Modeling

Simulations based on large language models (LLMs) have proven to be powerful for understanding human behavior, making them valuable additions to the social scientific toolkit. However, LLMs are ultimately black boxes based on deep neural networks which limits their value for social science. This is because of a lack of (i) interpretability: i.e. the ability to assign clear mechanisms driving observed behavior; and a lack of (ii) steerability: i.e. the ability to mute or amplify specific theoretically meaningful mechanisms of action to drive specific model behavior. Here, we demonstrate how the black box could be opened up to further enrich LLM-based simulations. Specifically, we compare three types of methods: (1) prompt-based manipulation, (2) SAE-derived feature steering, and (3) probe-based direction steering and examine their utility for LLM-based social scientific simulations. We do so by interpreting and steering two foundational components of human behaviors, namely preferences (risk attitudes, altruism) and capabilities (divergent creativity, product innovation), operationalized using four classic economic and creative tasks implemented as natural-language interactions.

Introduction. Large Language Models (LLMs) mimic humans on multiple dimensions: they understand natural Having said that, LLM-based simulations might not always faithfully match human behavior Here, we develop a framework to compare alternate methods for looking into the LLM black-box

Discussion / Conclusion. Future work might also explore hybrid approaches, such as multi-probe steering or low-rank

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Do language models reason like humans or mimic surface patterns? Why do persona simulations fail to predict authentic user behavior? How do false presuppositions and sycophancy drive persistent false beliefs in models? What design and behavioral factors drive false consciousness attribution to AI? Why doesn't reasoning volume improve theory of mind performance? What mechanisms preserve shared understanding in evolving conversations? What factors drive AI persuasiveness and how can it be mitigated? Where and how do personality traits reside in language models? Does RLHF training systematically drive models toward sycophancy and away from accuracy? How do agent-learned skills transfer and improve across different tasks? Do language models develop actual world models or merely task heuristics? Do language models reason through causal mechanisms or semantic associations? How well do AI systems understand human social norms?