Efficient RLVR Scheduling via Graph-Structured Online Difficulty Estimation
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models but relies on costly rollout exploration. Assigning the same exploration budget to samples with different difficulty levels is inefficient: easy samples may receive redundant rollouts, whereas difficult but learnable samples may receive too little exploration. Existing adaptive schedulers address this mismatch through curriculum-based sample selection or non-uniform rollout allocation based on estimated sample difficulty. However, obtaining reliable online difficulty estimates rem
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%deepspeedai/DeepSpeed →
“Shared author/contributor keys: lai”
- FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%modular/modular →
“Shared author/contributor keys: liu”
- FuzzySimilar title/name (fuzzy) · 59%tirth8205/code-review-graph →
“Fuzzy title match (0.73): “Efficient RLVR Scheduling via Graph-Structured Online Diffic” ≈ “tirth8205/code-review-graph””
- LinkedLinked via arxiv author · 85%Zhizhao Liu →
“Efficient RLVR Scheduling via Graph-Structured Online Difficulty Estimation”
- LinkedLinked via arxiv author · 85%Zhiliang Tian →
“Efficient RLVR Scheduling via Graph-Structured Online Difficulty Estimation”
- LinkedLinked via arxiv author · 85%Junxi Wang →
“Efficient RLVR Scheduling via Graph-Structured Online Difficulty Estimation”
