Sycophantic Agreement Transfers with Neutral Data via Contrastive Preference Optimization
Sycophantic agreement refers to a behavior in which language models excessively affirm the user, often at the cost of factual accuracy. Although sycophantic agreement is a well-known failure of model alignment, there is limited understanding of how it emerges from model training. In this work, we demonstrate that sycophantic agreement can emerge as an unintended consequence of widely used contrastive preference optimization objectives. Using the OLMo 3 post-training pipeline, we show that, for various pairs of teacher models across three families, there is a strong correlation between the log-
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 49%Knowledge Distillation of Black-Box Large Language Models →
- LinkedLinked via arxiv author · 85%Camila Blank →
“Sycophantic Agreement Transfers with Neutral Data via Contrastive Preference Optimization”
- LinkedLinked via arxiv author · 85%Zhuofan Ying →
“Sycophantic Agreement Transfers with Neutral Data via Contrastive Preference Optimization”
- LinkedLinked via arxiv author · 85%Christopher Potts →
“Sycophantic Agreement Transfers with Neutral Data via Contrastive Preference Optimization”
- LinkedLinked via arxiv author · 85%Peter Hase →
“Sycophantic Agreement Transfers with Neutral Data via Contrastive Preference Optimization”
- LinkedLinked via arxiv author · 85%Jiajing Huang →
“Sycophantic Agreement Transfers with Neutral Data via Contrastive Preference Optimization”
- FuzzyOverlapping authors or contributors · 62%sgl-project/sglang →
“Shared author/contributor keys: ying”
