How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?
Modern LLMs are alarmingly susceptible to surprisingly simple immaterial changes of input prompts: a casual hint, an incorrectly labeled few-shot example, or a fake prior assistant turn often flips an originally correct answer. We study where this susceptibility, spanning sycophancy and related cue-induced biases, lives inside the model. Across five model families and seven BCT bias types, we extract a per-bias direction from hidden states and triangulate it through three measures: probing, leave-one-dataset-out transfer, and causal intervention. The susceptibility is largely installed by alig
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 47%Alignment →
- PossiblePossibly related (embedding) · 45%When I made LLMs argue with each other, they started making up citations to win. Sycophancy wasn't the only failure mode. →
- FuzzyOverlapping authors or contributors · 62%microsoft/ML-For-Beginners →
“Shared author/contributor keys: gupta”
- FuzzyOverlapping authors or contributors · 62%keras-team/keras →
“Shared author/contributor keys: jin”
- FuzzyOverlapping authors or contributors · 62%HKUDS/LightRAG →
“Shared author/contributor keys: jin”
- LinkedLinked via arxiv author · 85%Prakhar Gupta →
“How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?”
- LinkedLinked via arxiv author · 85%Terry Jingchen Zhang →
“How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?”
- LinkedLinked via arxiv author · 85%Florent Draye →
“How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?”
