newsReddit r/MachineLearningTrust 52 · CommunityPublished 1mo agoLive · 1mo ago
New LLM Coordination Benchmark - Benchmarking Open-Ended Multi-Agent Coordination in Language Agents [R]
Can LLM agents coordinate in long-horizon, open-ended worlds? We evaluate 13 modern LLMs in a new benchmark where agents must work together to explore, communicate, trade resources, craft tools, build structures, and fight mobs. TL;DR : Most agents struggle, averaging only ~6% normalised return. Yet on the hardest setting, zero-shot Gemini 3.1 Pro performs comparably to the best MARL agent trained for 1 billion envi
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 58%Giskard-AI/giskard-oss →
- PossiblePossibly related (embedding) · 58%PACE: A Proxy for Agentic Capability Evaluation →
- PossiblePossibly related (embedding) · 56%Productive-Superintelligence/lllm →
- PossiblePossibly related (embedding) · 56%MetaSkill-Evolve: Recursive Self-Improvement of LLM Agents via Two-Timescale Meta-Skill Evolution →
- PossiblePossibly related (embedding) · 56%agent-tools →
- PossiblePossibly related (embedding) · 48%Superior-Trade/superior-skills →
- PossiblePossibly related (embedding) · 51%agentscope-ai/PawBench →
- PossiblePossibly related (embedding) · 56%Adaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent Security →
Covers
Covers (incoming)
repoSuperior-Trade/superior-skillsrepoagentscope-ai/PawBenchpaperAdaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent SecurityrepoAgentR1/Agent-R1repomskazemi/aobenchreporybawaw/TradingAgentsPLrepomipkovich/tame-swarmpaperThe Interaction Tax: When Communication Erases Diversity in Multi-Agent Teams
Related across the graph
repoProductive-Superintelligence/lllmrepomskazemi/aobenchrepoGiskard-AI/giskard-osspaperThe Interaction Tax: When Communication Erases Diversity in Multi-Agent TeamspaperMetaSkill-Evolve: Recursive Self-Improvement of LLM Agents via Two-Timescale Meta-Skill Evolutionreporybawaw/TradingAgentsPLrepoSuperior-Trade/superior-skillspaperPACE: A Proxy for Agentic Capability EvaluationrepoAgentR1/Agent-R1paperAdaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent Securityrepoagentscope-ai/PawBenchrepoagent-toolsrepomipkovich/tame-swarm
