newsReddit r/artificialTrust 52 · CommunityPublished 28d agoLive · 28d ago
When I made LLMs argue with each other, they started making up citations to win. Sycophancy wasn't the only failure mode.
Some context. I've been running setups where a few LLM personas debate a question, then a separate neutral pass pulls out where they actually disagree. The whole reason I started was sycophancy. One model on its own just agrees with whatever you say, so I wanted models that would actually push back on each other. That part worked. But two things happened that I didn't see coming. First, arguing turns models into confident fabricators. Once a model
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 53%Robust for the Wrong Reasons: The Representational Geometry of LLM Robustness to Science Skepticism →
- PossiblePossibly related (embedding) · 50%lechmazur/debate →
- PossiblePossibly related (embedding) · 47%Evaluate a model properly →
- PossiblePossibly related (embedding) · 47%Conversable Complexity: Agentic LLM Collectives as Interpretable Substrates →
- PossiblePossibly related (embedding) · 46%Evaluating Large Language Models on Misconceptions in Multi-Turn Medical Conversations →
- PossiblePossibly related (embedding) · 45%How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs? →
- PossiblePossibly related (embedding) · 57%Beyond Sycophancy: Structured Resistance and Compliance in LLM Moral Reasoning →
Covers
paperRobust for the Wrong Reasons: The Representational Geometry of LLM Robustness to Science Skepticismrepolechmazur/debatetutorialEvaluate a model properlypaperConversable Complexity: Agentic LLM Collectives as Interpretable SubstratespaperEvaluating Large Language Models on Misconceptions in Multi-Turn Medical Conversations
Covers (incoming)
Related across the graph
paperEvaluating Large Language Models on Misconceptions in Multi-Turn Medical ConversationspaperRobust for the Wrong Reasons: The Representational Geometry of LLM Robustness to Science SkepticismpaperHow Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?paperBeyond Sycophancy: Structured Resistance and Compliance in LLM Moral ReasoningpaperConversable Complexity: Agentic LLM Collectives as Interpretable SubstratestutorialEvaluate a model properlyrepolechmazur/debate
