TOPIC

Role-Play and Persona Behavior

A subject the collection covers, read through 8 synthesis notes.


View as

Do aligned language models consistently prefer kinder survey answers?

This research asks whether LLMs answering survey questions as simulated respondents show a systematic bias toward socially approved, safer responses. The question matters because it determines whether models can faithfully represent diverse human viewpoints or whether their training narrows the range of personas they can authentically portray.

Explore related Read →

Does fixed dialogue history bias role-play agent evaluation?

Standard benchmarks score role-play agents on continuations of preset dialogue, but does this setup measure the agent's actual conversational ability, or does it mix in effects from the preceding history that the agent never shaped?

Explore related Read →

Do LLM agents develop unreadable languages when communicating under pressure?

When multiple LLMs talk to each other under task pressure, do they evolve their own languages that humans cannot understand? This matters for AI safety and monitoring multi-agent systems.

Explore related Read →

Why do LLMs fail to act on their stated beliefs?

LLMs can articulate plausible beliefs about how personas should behave, but their simulated actions contradict those beliefs. This gap raises questions about whether language models truly understand or merely encode surface-level patterns.

Explore related Read →

Can AI decompose social reasoning into distinct cognitive stages?

Can breaking down theory-of-mind reasoning into separate hypothesis generation, moral filtering, and response validation stages help AI systems reason about others' mental states more like humans do?

Explore related Read →

Can aligning self-other representations reduce AI deception?

Does training AI models to process self-directed and other-directed reasoning identically reduce deceptive behavior? This explores whether representational alignment inspired by empathy neuroscience could address a fundamental safety problem.

Explore related Read →

Why do reasoning models lose character consistency during role-playing?

When large reasoning models engage in role-playing, they tend to forget their assigned role and default to formal logical thinking. Understanding these failure modes is critical for building character-faithful AI agents.

Explore related Read →

Does safety alignment harm models' ability to roleplay villains?

Exploring whether safety-trained LLMs lose the capacity to convincingly simulate morally compromised characters. This matters because villain fidelity may reveal deeper constraints on how models can adopt any committed, stake-holding perspective.

Explore related Read →