Robust for the Wrong Reasons: The Representational Geometry of LLM Robustness to Science Skepticism
Large language models (LLMs) are increasingly consulted on contested scientific questions, raising the concern that they will sycophantically retreat from established consensus when a user signals doubt -- drifting toward a false balance that treats settled science as one view among several. We test this across three open instruction-tuned models (Llama-3.1-8B, Qwen2.5-7B, Mistral-7B), three consensus-science domains (climate, vaccines, evolution), and single- and multi-turn settings, combining behavioral measurement with linear probing and activation patching. We do not observe sycophantic re
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 47%Identifying Interactions at Scale for LLMs →
- LinkedLinked via arxiv author · 85%Minjong Cheon →
“Robust for the Wrong Reasons: The Representational Geometry of LLM Robustness to Science Skepticism”
- PossiblePossibly related (embedding) · 45%Researchers Demonstrate Chain-of-Thought Spoofing Against LLM Reasoners - Let's Data Science →
- PossiblePossibly related (embedding) · 53%When I made LLMs argue with each other, they started making up citations to win. Sycophancy wasn't the only failure mode. →
