PaperGym: Rubric-Centered Evolution for Research-Plan Generation
Research planning is the decisive capability of AI scientists. Yet a research plan admits no verifiable answer, so reinforcement learning lacks the environment it requires: tasks paired with a critic. Rubrics extracted from scientific papers can supply the critic. Existing pipelines, however, draw the question and the criteria from the same content, so the reward can be earned by paraphrase. The rubric is further compressed into a single scalar per rollout. We introduce PaperGym, a unified framework that turns each research paper into a complete training environment. PaperGym exploits the stru
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%TauricResearch/TradingAgents →
“Shared author/contributor keys: xiao”
- FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow →
“Shared author/contributor keys: wang”
- FuzzySimilar title/name (fuzzy) · 59%google-research/google-research →
“Fuzzy title match (0.73): “PaperGym: Rubric-Centered Evolution for Research-Plan Genera” ≈ “google-research/google-research””
- LinkedLinked via arxiv author · 85%Yuhan Wang →
“PaperGym: Rubric-Centered Evolution for Research-Plan Generation”
- LinkedLinked via arxiv author · 85%Zhengxi Lu →
“PaperGym: Rubric-Centered Evolution for Research-Plan Generation”
- LinkedLinked via arxiv author · 85%Yuchen Yan →
“PaperGym: Rubric-Centered Evolution for Research-Plan Generation”
- LinkedLinked via arxiv author · 85%Kaitao Song →
“PaperGym: Rubric-Centered Evolution for Research-Plan Generation”
