Read original ↗
paperarXivTrust 82 · PrimaryPublished 6d agoLive · 5d ago

SPADE: Self-Play in Adaptive Synthetic Executable Environments

Continuous self-improvement requires an ever-expanding pool of self-generated, diverse, adaptive goals. For language agents, existing training environment pools (hand-curated, statically synthesized, or frozen-verifier) keep the goal distribution fixed as the learner scales. We introduce SPADE (Self-Play in Adaptive Synthetic Executable Environments), a self-play RL framework in which a single LLM plays two roles: an Environment Designer that writes complete, long-horizon training environments as executable code with an OpenAI Gym-style reset()/step() interface, and a Reasoning Agent that lear

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzyOverlapping authors or contributors · 62%affaan-m/ECC

    Shared author/contributor keys: jiang

  • FuzzyOverlapping authors or contributors · 62%modular/modular

    Shared author/contributor keys: liu

  • FuzzyOverlapping authors or contributors · 62%ultralytics/yolov5

    Shared author/contributor keys: choi

  • FuzzyOverlapping authors or contributors · 62%BerriAI/litellm

    Shared author/contributor keys: jiang

  • FuzzyOverlapping authors or contributors · 62%sgl-project/sglang

    Shared author/contributor keys: zhou

  • LinkedLinked via arxiv author · 85%Bo Liu

    SPADE: Self-Play in Adaptive Synthetic Executable Environments

  • LinkedLinked via arxiv author · 85%Simon Yu

    SPADE: Self-Play in Adaptive Synthetic Executable Environments

  • LinkedLinked via arxiv author · 85%Yiding Jiang

    SPADE: Self-Play in Adaptive Synthetic Executable Environments

Implements (incoming)

authored (incoming)

Related across the graph

Topics