Read original ↗
paperarXivTrust 82 · PrimaryPublished 29d agoLive · 25d ago

Understanding Reasoning from Pretraining to Post-Training

Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning tasks, yet RL post-training is largely studied in isolation from the pretraining that precedes it. As a result, two basic questions remain open: (1) how do pretraining choices (model size, data) shape the returns to RL compute, and (2) what does RL actually do to the model? These questions are difficult to study in the standard LLM setting: pretraining corpora are vast and uncontrolled, making it hard to attribute behaviors to pretraining versus RL, and systematic compute sweeps across

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • PossiblePossibly related (embedding) · 54%RLHF
  • FuzzyOverlapping authors or contributors · 62%google-research/google-research

    Shared author/contributor keys: sun

  • LinkedLinked via arxiv author · 85%Jingyan Shen

    Understanding Reasoning from Pretraining to Post-Training

  • LinkedLinked via arxiv author · 85%Chenyang Li

    Understanding Reasoning from Pretraining to Post-Training

  • LinkedLinked via arxiv author · 85%Salman Rahman

    Understanding Reasoning from Pretraining to Post-Training

  • LinkedLinked via arxiv author · 85%Yifan Sun

    Understanding Reasoning from Pretraining to Post-Training

  • LinkedLinked via arxiv author · 85%Micah Goldblum

    Understanding Reasoning from Pretraining to Post-Training

  • LinkedLinked via arxiv author · 85%Matus Telgarsky

    Understanding Reasoning from Pretraining to Post-Training

Related to

Implements (incoming)

authored (incoming)

Related across the graph

Topics