repoGitHubTrust 82 · PrimaryPublished 29d agoLive · 21d ago
lechmazur/debate
Adversarial multi-turn benchmark for LLM debate quality, using side-swapped matchups and multi-model judging to rank models by judged debate performance.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 50%Evaluate a model properly →
- PossiblePossibly related (embedding) · 50%When I made LLMs argue with each other, they started making up citations to win. Sycophancy wasn't the only failure mode. →
