Read original ↗
paperarXivTrust 82 · PrimaryPublished 29d agoLive · 28d ago

PPL-Factory: Task-Aware and Budget-Aware Data Selection from Language Modeling to Reasoning

Not all training samples contribute equally to large language model fine-tuning. Selecting informative training samples can reduce the computational cost while preserving downstream performance. Many existing data selection methods rely on indirect heuristics, such as data quality, diversity or reasoning trace length. However, the effectiveness of these fixed criteria is task-dependent and difficult to generalize across diverse downstream tasks. Perplexity-based data selection provides a simple and model-aware solution to estimate the sample difficulty, but existing approaches typically score

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzySimilar title/name (fuzzy) · 59%hiyouga/LlamaFactory

    Fuzzy title match (0.73): “PPL-Factory: Task-Aware and Budget-Aware Data Selection from” ≈ “hiyouga/LlamaFactory”

  • LinkedLinked via arxiv author · 85%Yuhang Zhang

    PPL-Factory: Task-Aware and Budget-Aware Data Selection from Language Modeling to Reasoning

  • LinkedLinked via arxiv author · 85%Warren J. Gross

    PPL-Factory: Task-Aware and Budget-Aware Data Selection from Language Modeling to Reasoning

Implements (incoming)

authored (incoming)

Related across the graph

Topics