newsReddit r/MachineLearningTrust 72 · CommunityPublished 1mo agoLive · 1mo ago
I made a superhuman Generals.io agent with self-play RL [P]
Hi everyone, I trained a self-play RL agent for Generals.io that reached superhuman-level and ranked #1 on the human 1v1 leaderboard. It began as my master's thesis where the goal was to beat a prior algorithm based agent. We succeeded using behavior cloning, RL fine-tuning and reward shaping, but the agent was still consistently beaten by the top players. So I gave it a round two and fixed the largest bottle
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownSelf-rewarding agents that retrace failures →
- LinkedLinked via unknownretrace-agents →
- LinkedLinked via unknownagent-tools →
- LinkedLinked via unknownNoshkoto/Noshy →
- LinkedLinked via unknownkuzmenkoff/exec_magica_engine →
- LinkedLinked via unknownTriadic Werewolf: A Jester Role for Multi-Hop Theory of Mind in LLMs →
- LinkedLinked via unknownjaimasih05-commits/swarm-foraging-qlearn →
- PossiblePossibly related (embedding) · 47%tomasz-tomczyk/crit →
Covers
Covers (incoming)
repoNoshkoto/Noshyrepokuzmenkoff/exec_magica_enginepaperTriadic Werewolf: A Jester Role for Multi-Hop Theory of Mind in LLMsrepojaimasih05-commits/swarm-foraging-qlearnrepotomasz-tomczyk/critreposteveyeow/FeynmanrepoShiyao-Huang/awesome-agent-evolutionrepoAgentToolkit/altk-evolvepaperTeach it to stop, not just to clickrepoAlibaba-Quark/SSP
Related across the graph
repotomasz-tomczyk/critpaperTriadic Werewolf: A Jester Role for Multi-Hop Theory of Mind in LLMsrepoShiyao-Huang/awesome-agent-evolutionrepojaimasih05-commits/swarm-foraging-qlearnpaperSelf-rewarding agents that retrace failuresrepoAgentToolkit/altk-evolverepokuzmenkoff/exec_magica_enginereporetrace-agentsreposteveyeow/FeynmanrepoAlibaba-Quark/SSPrepoNoshkoto/Noshyrepoagent-toolspaperTeach it to stop, not just to click
