Read original ↗
paperarXivTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago

Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape Pre-, Intra-, and Post-CoT Calibration

Large language models have made strong reasoning gains through supervised fine-tuning, reinforcement learning, and on-policy distillation, yet these post-training methods are usually evaluated only by final-answer accuracy. We study how they reshape confidence during reasoning. We introduce a three-stage calibration framework that evaluates confidence before, during, and after chain-of-thought generation, corresponding to difficulty estimation, early termination, and answer aggregation. Through a controlled comparison on mathematical reasoning benchmarks, we find that OPD provides the most use

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • PossiblePossibly related (embedding) · 53%New benchmark exposes reasoning gaps in top models
  • FuzzyOverlapping authors or contributors · 62%hiyouga/LlamaFactory

    Shared author/contributor keys: lin

  • FuzzyOverlapping authors or contributors · 62%Zeyi-Lin/HivisionIDPhotos

    Shared author/contributor keys: lin

  • LinkedLinked via arxiv author · 85%Shuhao Li

    Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape Pre-, Intra-, and Post-CoT Calibra

  • LinkedLinked via arxiv author · 85%Guodong Du

    Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape Pre-, Intra-, and Post-CoT Calibra

  • LinkedLinked via arxiv author · 85%Anhao Zhao

    Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape Pre-, Intra-, and Post-CoT Calibra

  • LinkedLinked via arxiv author · 85%Wanyu Lin

    Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape Pre-, Intra-, and Post-CoT Calibra

  • LinkedLinked via arxiv author · 85%Tianyu Yuan

    Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape Pre-, Intra-, and Post-CoT Calibra

Covers

Implements (incoming)

authored (incoming)

Related across the graph

Topics