Gotta Catch them all: the modes of Sycophancy
Large language models often align with users' beliefs at the expense of factual accuracy, a behavior known as sycophancy. Prior mechanistic studies largely treat sycophancy as a single behavioral dimension that can be uniformly amplified or suppressed. We challenge this assumption by analyzing three hypothesized modes of sycophancy across 948 social pressure situations. Although the modes produce highly similar outputs, with a text-only classifier achieving just 57.8 percent accuracy, their internal representations are perfectly linearly separable from layer 14 onward. We further find the mode
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 56%Understanding large language models demands distinguishing human projection from machine cognition - Nature →
- PossiblePossibly related (embedding) · 51%Evaluating J-space entropy as an error predictor across 7 datasets on Qwen3-4B [R] →
- PossiblePossibly related (embedding) · 50%Benchmarking large language models against practicing clinicians on psychopathological assessment - Nature →
- PossiblePossibly related (embedding) · 49%Large language models exhibit stigmatizing behaviour in contextual judgements of health conditions - Nature →
- FuzzyOverlapping authors or contributors · 62%firecrawl/firecrawl →
“Shared author/contributor keys: jain”
- LinkedLinked via arxiv author · 85%Shreyans Jain →
“Gotta Catch them all: the modes of Sycophancy”
- LinkedLinked via arxiv author · 85%Alexandra Yost →
“Gotta Catch them all: the modes of Sycophancy”
- LinkedLinked via arxiv author · 85%Amirali Abdullah →
“Gotta Catch them all: the modes of Sycophancy”
