SYNTHESIS NOTE
Topics›Recommenders LLMs›this note

Do comparisons help users evaluate items better than isolated descriptions?

Can framing product evaluations relationally—by comparing to other items—ground assessment in user reasoning better than absolute descriptions? This matters because recommendation explanations often ask users to do comparison work mentally.

Synthesis note · 2026-05-03 · sourced from Recommenders LLMs

Standard recommendation explanations evaluate items in isolation: "this piano sounds natural." A user has to do the comparison work in their head, judging this evaluation against their experience with other pianos. Comparative recommendations ground the evaluation by referencing another item: "This piano sounds more natural than my Sony NWZ-A855." The relational frame embeds the comparison the user would otherwise construct.

Comparing Apples to Apples generates these comparative sentences from user reviews. A BERT classifier, fine-tuned on manually labeled examples, identifies comparative sentences in product reviews. From a corpus of 258,816 comparative sentences and associated reviews, the system extracts aspects (sound quality, price-to-value, longevity) and their associated sentiments per item. These aspects feed into abstractive generation: the system generates new comparative sentences highlighting features relevant to a particular user, using product and user information as conditioning.

Two aspects are personalizable: which features matter to the user (extracted from their review history), and which positive or negative aspects to emphasize. A user who has historically focused on price will get price comparisons; one who has focused on sound quality will get sound comparisons. Human evaluation on Comparativeness, Relevance, and Fidelity confirms the generated sentences are both true to the source material and useful for purchase decisions.

The general principle: when evaluation is the goal, relational explanations carry more information than absolute ones because relational framing matches how humans evaluate. A recommendation system producing relational descriptions is closer to user reasoning than one that lists attributes per item.

Inquiring lines that read this note 12

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do social dynamics distort aggregated online ratings? How does the generation-verification gap limit what we can measure about AI reasoning? How do false presuppositions and sycophancy drive persistent false beliefs in models? Can preference-based training achieve better behavior optimization than supervised fine-tuning alone? Why do some clarifying approaches produce understanding while others just satisfy? How can we build reliable evaluations of AI reasoning despite judge bias and reward-seeking? What factors drive AI persuasiveness and how can it be mitigated?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
16 direct connections · 92 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

comparative recommendations ground item evaluation by referencing other items — abstractive aspect-controlled generation from review-extracted aspects