Rating the Pitch, Not the Product: User Evaluations of LLMs Reflect Expectations More Than Performance
Imagine two users interact with the same LLM. One has been told it is the cutting-edge flagship model; the other, an older, weaker model. They walk away with markedly different ratings of its usefulness and intelligence, yet they used the same model. In a controlled study, 162 participants each used one of six LLMs from two families across three collaborative tasks, after first viewing a landing page that matched, overstated, or understated their model's true capability. This pre-interaction framing shifted user opinions and interaction behavior while task performance did not. Oversold users r
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 49%Evaluate a model properly →
- LinkedLinked via arxiv author · 85%Robert Morabito →
“Rating the Pitch, Not the Product: User Evaluations of LLMs Reflect Expectations More Than Performance”
- LinkedLinked via arxiv author · 85%Tyler McDonald →
“Rating the Pitch, Not the Product: User Evaluations of LLMs Reflect Expectations More Than Performance”
- LinkedLinked via arxiv author · 85%Charitra Viswanath →
“Rating the Pitch, Not the Product: User Evaluations of LLMs Reflect Expectations More Than Performance”
- LinkedLinked via arxiv author · 85%Angel Hsing-Chi Hwang →
“Rating the Pitch, Not the Product: User Evaluations of LLMs Reflect Expectations More Than Performance”
- LinkedLinked via arxiv author · 85%Susanne Gaube →
“Rating the Pitch, Not the Product: User Evaluations of LLMs Reflect Expectations More Than Performance”
- LinkedLinked via arxiv author · 85%Jad Kabbara →
“Rating the Pitch, Not the Product: User Evaluations of LLMs Reflect Expectations More Than Performance”
- LinkedLinked via arxiv author · 85%Ali Emami →
“Rating the Pitch, Not the Product: User Evaluations of LLMs Reflect Expectations More Than Performance”
