Read original ↗
paperarXivTrust 82 · PrimaryPublished 10d agoLive · 9d ago

S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?

Large language models (LLMs) increasingly interact with external environments and accumulate substantial behavioral experience, yet existing agent benchmarks largely evaluate them as fixed policies. It therefore remains unclear whether an agent can actively test its behavior, judge the resulting experience, and use that experience to improve future decisions. We introduce \textbf{S\textsuperscript{3}Gym}, an interactive benchmark for evaluating LLM self-improvement through three coupled capabilities: \textbf{Self-Testing}, \textbf{Self-Judging}, and \textbf{Self-Improvement}. S$^3$Gym separate

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • PossiblePossibly related (embedding) · 59%Can AI Improve Itself? RSI Might Be the Answer [R]
  • FuzzyOverlapping authors or contributors · 62%modular/modular

    Shared author/contributor keys: liu

  • FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow

    Shared author/contributor keys: wang

  • FuzzyOverlapping authors or contributors · 62%ray-project/ray

    Shared author/contributor keys: wang

  • LinkedLinked via arxiv author · 85%Jiajun Shi

    S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?

  • LinkedLinked via arxiv author · 85%Siyuan Tao

    S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?

  • LinkedLinked via arxiv author · 85%Yuhao Wu

    S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?

  • LinkedLinked via arxiv author · 85%Zexuan Wang

    S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?

Covers

Implements (incoming)

authored (incoming)

Related across the graph

Topics