S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?
Large language models (LLMs) increasingly interact with external environments and accumulate substantial behavioral experience, yet existing agent benchmarks largely evaluate them as fixed policies. It therefore remains unclear whether an agent can actively test its behavior, judge the resulting experience, and use that experience to improve future decisions. We introduce \textbf{S\textsuperscript{3}Gym}, an interactive benchmark for evaluating LLM self-improvement through three coupled capabilities: \textbf{Self-Testing}, \textbf{Self-Judging}, and \textbf{Self-Improvement}. S$^3$Gym separate
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 59%Can AI Improve Itself? RSI Might Be the Answer [R] →
- FuzzyOverlapping authors or contributors · 62%modular/modular →
“Shared author/contributor keys: liu”
- FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- LinkedLinked via arxiv author · 85%Jiajun Shi →
“S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?”
- LinkedLinked via arxiv author · 85%Siyuan Tao →
“S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?”
- LinkedLinked via arxiv author · 85%Yuhao Wu →
“S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?”
- LinkedLinked via arxiv author · 85%Zexuan Wang →
“S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?”
