Read original ↗
paperarXivTrust 82 · PrimaryPublished 2d agoLive · 3h ago

Envs-FORGE: Frontier-Optimized Reward-Grounded Environment Synthesis for Agent RL

Reinforcement learning (RL) for terminal agents needs executable training environments with reliable rewards and useful difficulty. Fixed recipes such as few-shot, Self-Instruct, and Evol-Instruct apply the same prompting policy to every seed, even when the current policy would benefit from a harder, easier, or simply different task. We present Envs-FORGE, a prompting policy that converts verifier rewards into per-seed environment-synthesis actions. Envs-FORGE estimates seed pass rates, scores six projection--direction actions around a target learning frontier, and solves a per-seed mixed-inte

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzySimilar title/name (fuzzy) · 87%CodeForge-15B

    Fuzzy title match (0.94): “Envs-FORGE: Frontier-Optimized Reward-Grounded Environment S” ≈ “CodeForge-15B”

  • FuzzySimilar title/name (fuzzy) · 59%AgentCore-8B

    Fuzzy title match (0.73): “Envs-FORGE: Frontier-Optimized Reward-Grounded Environment S” ≈ “AgentCore-8B”

  • FuzzySimilar title/name (fuzzy) · 87%SWE-agent/SWE-agent

    Fuzzy title match (0.94): “Envs-FORGE: Frontier-Optimized Reward-Grounded Environment S” ≈ “SWE-agent/SWE-agent”

  • FuzzySimilar title/name (fuzzy) · 87%zhayujie/CowAgent

    Fuzzy title match (0.94): “Envs-FORGE: Frontier-Optimized Reward-Grounded Environment S” ≈ “zhayujie/CowAgent”

  • FuzzyOverlapping authors or contributors · 62%affaan-m/ECC

    Shared author/contributor keys: jiang

  • FuzzyOverlapping authors or contributors · 62%BerriAI/litellm

    Shared author/contributor keys: jiang

  • FuzzyOverlapping authors or contributors · 62%Zeyi-Lin/HivisionIDPhotos

    Shared author/contributor keys: lin

  • LinkedLinked via arxiv author · 85%Xiaojun Wu

    Envs-FORGE: Frontier-Optimized Reward-Grounded Environment Synthesis for Agent RL

Has model

Implements (incoming)

authored (incoming)

Related across the graph

Topics