Read original ↗
newsReddit r/artificialTrust 52 · CommunityPublished 28d agoLive · 28d ago

When I made LLMs argue with each other, they started making up citations to win. Sycophancy wasn't the only failure mode.

Some context. I've been running setups where a few LLM personas debate a question, then a separate neutral pass pulls out where they actually disagree. The whole reason I started was sycophancy. One model on its own just agrees with whatever you say, so I wanted models that would actually push back on each other. That part worked. But two things happened that I didn't see coming. First, arguing turns models into confident fabricators. Once a model

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

Covers

Covers (incoming)

Related across the graph