Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence
Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR) can substantially improve reasoning in mathematics and code, where outcomes can be checked automatically. Extending this progress to open-ended and agentic tasks remains difficult because reliable rewards are harder to obtain and direct human supervision cannot keep pace with the scale and complexity of model-generated experience. This paper studies how LRMs can continue to improve as human supervision gradually recedes from the learning loop. We examine two connected dimensio
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 84%roboflow/supervision →
“Fuzzy title match (0.92): “Scaling Large Reasoning Models beyond Human Supervision: A P” ≈ “roboflow/supervision””
- FuzzyOverlapping authors or contributors · 62%janhq/jan →
“Shared author/contributor keys: han”
- FuzzyOverlapping authors or contributors · 62%Zeyi-Lin/HivisionIDPhotos →
“Shared author/contributor keys: lin”
- FuzzyOverlapping authors or contributors · 62%ultralytics/ultralytics →
“Shared author/contributor keys: han”
- FuzzyOverlapping authors or contributors · 62%google-research/google-research →
“Shared author/contributor keys: sun”
- LinkedLinked via arxiv author · 85%Zhiqin Yang →
“Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence”
- LinkedLinked via arxiv author · 85%Jingwen Fu →
“Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence”
- LinkedLinked via arxiv author · 85%Yuhan Liu →
“Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence”
