SYNTHESIS NOTE
Topics›Personalization›this note

Do user outputs outperform inputs for LLM personalization?

Does a user's history of outputs (responses, endorsed content) matter more for personalization than their input queries? This explores what actually drives effective personalization in language models.

Synthesis note · 2026-02-23 · sourced from Personalization

A study on user profile roles in LLM personalization surfaces a counterintuitive finding: the outputs users have produced or endorsed matter far more than the inputs they submitted. Using only the output part of user profiles achieves comparable or even superior performance to complete profiles across multiple LaMP tasks. Using only the input part leads to noticeable degradation.

This finding separates personalization from two adjacent paradigms:

Personalization ≠ RAG. Retrieval-augmented generation relies on semantic similarity between the input query and retrieved documents. Personalization works through a different mechanism — it is the style, preferences, and judgments expressed in historical responses that calibrate the model, not the semantic content of past queries.

Personalization ≠ ICL. In-context learning uses complete input-output pairs as demonstrations. Personalization requires only the output side — the response patterns that reveal who the user is and what they value.

The practical implication: when designing personalization systems under input length constraints, prioritize incorporating user-generated or user-approved responses over query histories. This unlocks the potential to include many more user profiles within limited context windows, because output-only profiles are both more effective and more compact than complete interaction histories.

A secondary finding adds a structural dimension: user profiles integrated closer to the beginning of the input context have more influence on personalization than those placed elsewhere. This parallels the positional bias documented in ICL — since How much does demo position alone affect in-context learning accuracy?, the spatial attention pattern appears to be domain-general, affecting personalization placement decisions as well as few-shot learning.

The output-over-input finding connects to the broader question of what personalization is. Since Can text summaries beat embeddings for personalized reward models?, the PLUS approach of training a summarizer to extract preference dimensions rather than topic summaries from user history is vindicated — preference dimensions are properties of outputs, not inputs.

Inquiring lines that read this note 51

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Does abstract user knowledge outperform concrete interaction history in personalization? Do structural constraints outperform deep architectures in recommendation systems? How can persona-attention mechanisms improve both recommendation quality and explainability? What makes personas effective for predicting individual preferences and behavior? What factors drive AI persuasiveness and how can it be mitigated? How do LLM judges' systematic biases affect alignment and evaluation outcomes? Why do LLM recommenders underperform collaborative filtering despite their capabilities? What compositional reasoning failures limit large language models despite scale? Can language models build genuine grounding through interaction? What safeguards enable trustworthy AI-assisted scientific peer review at scale? When do multi-agent systems provide sufficient quality returns on token investment? What drives appropriate trust calibration in personalized AI systems? How do recommenders balance exploiting fresh signals against maintaining preference stability? How can reward models capture diverse human preferences without excluding minority populations? Does alignment training create genuine alignment or just output compliance? Can preference-based training achieve better behavior optimization than supervised fine-tuning alone? Is language model reasoning authentic and what causes models to reason? What capability trade-offs arise from domain specialization through fine-tuning? How does persona conditioning amplify demographic stereotyping and bias in models?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
18 direct connections · 116 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

historical user outputs drive personalization more effectively than input queries — personalization information not semantic information is the active ingredient