Understanding Reasoning from Pretraining to Post-Training
Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning tasks, yet RL post-training is largely studied in isolation from the pretraining that precedes it. As a result, two basic questions remain open: (1) how do pretraining choices (model size, data) shape the returns to RL compute, and (2) what does RL actually do to the model? These questions are difficult to study in the standard LLM setting: pretraining corpora are vast and uncontrolled, making it hard to attribute behaviors to pretraining versus RL, and systematic compute sweeps across
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 54%RLHF →
- FuzzyOverlapping authors or contributors · 62%google-research/google-research →
“Shared author/contributor keys: sun”
- LinkedLinked via arxiv author · 85%Jingyan Shen →
“Understanding Reasoning from Pretraining to Post-Training”
- LinkedLinked via arxiv author · 85%Chenyang Li →
“Understanding Reasoning from Pretraining to Post-Training”
- LinkedLinked via arxiv author · 85%Salman Rahman →
“Understanding Reasoning from Pretraining to Post-Training”
- LinkedLinked via arxiv author · 85%Yifan Sun →
“Understanding Reasoning from Pretraining to Post-Training”
- LinkedLinked via arxiv author · 85%Micah Goldblum →
“Understanding Reasoning from Pretraining to Post-Training”
- LinkedLinked via arxiv author · 85%Matus Telgarsky →
“Understanding Reasoning from Pretraining to Post-Training”
