REER-PT: Reverse-Engineered Reasoning for Perplexity-Guided Pre-training Data Augmentation
As language-model compute continues to scale, high-quality training data is becoming an increasingly important bottleneck. Conventional next-token prediction supervises what follows a context but leaves the intermediate reasoning behind that continuation implicit. We introduce \textbf{REER-PT}, a scalable framework that extends Reverse-Engineered Reasoning (REER) to raw pre-training data. REER-PT identifies continuations that are difficult to predict but can still be inferred from the preceding context, and inserts concise reasoning annotations that reconstruct the missing connection between c
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 53%Exploring self-distilled reasoning for supervised fine-tuning with Amazon Nova →
- FuzzyOverlapping authors or contributors · 62%modular/modular →
“Shared author/contributor keys: liu”
- LinkedLinked via arxiv author · 85%Haoran Que →
“REER-PT: Reverse-Engineered Reasoning for Perplexity-Guided Pre-training Data Augmentation”
- LinkedLinked via arxiv author · 85%Jiajun Shi →
“REER-PT: Reverse-Engineered Reasoning for Perplexity-Guided Pre-training Data Augmentation”
- LinkedLinked via arxiv author · 85%Ting Huang →
“REER-PT: Reverse-Engineered Reasoning for Perplexity-Guided Pre-training Data Augmentation”
- LinkedLinked via arxiv author · 85%Renming Pang →
“REER-PT: Reverse-Engineered Reasoning for Perplexity-Guided Pre-training Data Augmentation”
- LinkedLinked via arxiv author · 85%Jiaheng Liu →
“REER-PT: Reverse-Engineered Reasoning for Perplexity-Guided Pre-training Data Augmentation”
- LinkedLinked via arxiv author · 85%Ge Zhang →
“REER-PT: Reverse-Engineered Reasoning for Perplexity-Guided Pre-training Data Augmentation”
