Efficiently Estimating Optimal Hyperparameter Scaling Laws through Power-Law Entropy Search
Optimal hyperparameter scaling laws describe how the best hyperparameters for large language model (LLM) training change with model and data scale, enabling practitioners to predict optimal configurations at production scales without expensive large-scale tuning. However, estimating these scaling laws conventionally requires exhaustive grid searches over thousands of training runs, consuming enormous computational resources. We introduce Power-Law Entropy Search (PLES), a computational cost-aware acquisition function built on multi-fidelity Bayesian optimization that efficiently estimates opti
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 50%Run large language models with hundreds of billions of parameters locally! Apple’s new Mac packs hardware for AI, pushing prices to a new high. - Moomoo →
- PossiblePossibly related (embedding) · 49%Hyperparameters fine tuning for MARL comparative study [D] →
- PossiblePossibly related (embedding) · 48%If you had a 300M parameter model, what would you optimize it for? →
- PossiblePossibly related (embedding) · 47%Large language models as uncertainty-calibrated optimizers for experimental discovery →
- FuzzyOverlapping authors or contributors · 62%Zeyi-Lin/HivisionIDPhotos →
“Shared author/contributor keys: lin”
- FuzzyOverlapping authors or contributors · 62%hiyouga/LlamaFactory →
“Shared author/contributor keys: lin”
- LinkedLinked via arxiv author · 85%Zhiliang Chen →
“Efficiently Estimating Optimal Hyperparameter Scaling Laws through Power-Law Entropy Search”
- LinkedLinked via arxiv author · 85%Sebastian Ament →
“Efficiently Estimating Optimal Hyperparameter Scaling Laws through Power-Law Entropy Search”
