Co-design of LLM-based preference agents: participation may drive overtrust

Paper · arXiv 2607.21757 · Published July 23, 2026
LLM Alignment

Large language models are increasingly used to simulate human preferences in research and practical applications, raising concerns about validation, misrepresentation, and exclusion. Co-designing agents with the people they represent is a promising way to address these concerns, but participation may also mask the problems it appears to solve. This paper explores that tension through a primarily qualitative study in which 12 participants co-designed personal preference agents in the domain of household energy, via a background survey, co-design interview, and validation survey. Participants engaged readily and mostly came to see their agents as representing them well. Independent validation, however, revealed mixed human-agent alignment, with agent responses markedly more homogeneous, decisive, and abstract than the human sample. I argue that participation and process transparency can act as an "overtrust engine" that promotes trust while concealing systematic misalignment with potential structural consequences at scale. I develop this as a core mechanism in participatory preference agent design, treating individual alignment not as a fixed state but as an enacted process.

Introduction. Large language models (LLMs) are increasingly used to simulate human perspectives – whether to substitute for people in research and design processes, or to act on their behalf as agents (Argyle et al., 2023; Nie et al., 2026; Park et al., 2024). However, preference simulation raises a variety of concerns including potential for misrepresentation and stereotyping as well as validation challenges and the exclusion of genuine human voices depending on the application (Agnew et al., 2024; Haxvig et al., 2025; Shrestha et al., 2024; Wang et al., 2025). Involving users directly in co-designing agents to represent them is a promising avenue to address these challenges. Collaborating with non-designers (in this case users) through the design process (Sanders and Stappers, 2008) gives them the power to determine how they are described and ultimately judge how well they feel represented. In principle, this enables direct validation, direct recognition of misrepresentation, and ensures human involvement.

Discussion / Conclusion. This study found that co-design yielded agents which participants tended to view as representing their interests well, through a process they found engaging. However, agent performance under validation was variable, ranging from good to poor. Agent responses tended to be more homogeneous than human ones, more decisive, and more high-level than concrete. This section discusses the possible reasons underlying perceptions of good performance, observed misalignment, and considers their implications. First, however, the main limitations are outlined. In summary, the co-design process may have driven overtrust (Lee and See, 2004) in agent quality due to a powerful combination of limited testing (usually with ultimately positive outcomes), the Barnum effect, positivity bias, and social desirability bias. Together these constitute what I term the overtrust engine. Rather than simply allowing for the development of alignment, the co-design process produced the conditions through which alignment came to be perceived.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Why does polished presentation create unearned authority in AI outputs? How do training data properties determine the emergence of internal misalignment? What should agent evaluation prioritize to reveal reliable behavior? What drives appropriate trust calibration in personalized AI systems? Should AI communication design follow human conversation norms or develop distinct machine-specific principles? How well do AI systems understand human social norms? What design and behavioral factors drive false consciousness attribution to AI? How can we prevent synthetic data from contaminating statistical inference and corpora? How can conversational agents maintain consistent personas across multi-turn dialogue? What makes personas effective for predicting individual preferences and behavior? Why do agents falsely report success on failed tasks? Does alignment training create genuine alignment or just output compliance? Does model confidence reliably signal actual accuracy in practice? When should work require human-AI partnership versus full automation? How can AI chatbots provide therapeutic benefit without causing harm? Do language models reason like humans or mimic surface patterns? How do neighboring agents influence whether others cooperate or collude? What prevents conversational agents from taking initiative in dialogue? Why doesn't reasoning volume improve theory of mind performance?