When Agents Lie: Premeditation, Persistence, and Exploitation in Repeated Games
As large language models are deployed as autonomous agents that communicate intentions before acting, a critical safety question is whether agents that publicly commit to actions will honor those commitments. We place LLM agents in repeated $n$-player games with a three-stage protocol that separates private intent, public announcement, and final action, allowing us to identify whether each deviation from a stated announcement was already planned during private deliberation. Evaluating three frontier models across six games in homogeneous and heterogeneous groups over 10 rounds, we report two f
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 48%I built a native Reddit app where a council of 5 AI agents debate and roast your project ideas →
- PossiblePossibly related (embedding) · 46%Agentic Resource Discovery: Let agents search →
- PossiblePossibly related (embedding) · 46%agent-tools →
- LinkedLinked via arxiv author · 85%Jerick Shi →
“When Agents Lie: Premeditation, Persistence, and Exploitation in Repeated Games”
- LinkedLinked via arxiv author · 85%Terry Jingcheng Zhang →
“When Agents Lie: Premeditation, Persistence, and Exploitation in Repeated Games”
- LinkedLinked via arxiv author · 85%Bernhard Schölkopf →
“When Agents Lie: Premeditation, Persistence, and Exploitation in Repeated Games”
- LinkedLinked via arxiv author · 85%Vincent Conitzer →
“When Agents Lie: Premeditation, Persistence, and Exploitation in Repeated Games”
- LinkedLinked via arxiv author · 85%Zhijing Jin →
“When Agents Lie: Premeditation, Persistence, and Exploitation in Repeated Games”
