Super Weights in LLMs and the Failure of Selective Training
Recent work identified Super Weights, individual parameters whose removal degrades model performance by orders of magnitude. We show that this degradation due to pruning Super Weights does not universally apply to all LLMs. Furthermore, if these parameters are so important, Super Weight-aware training should be effective. We show the opposite. Training Super Weights in isolation (100 to 8,192 parameters) drops accuracy to random-guessing levels on both OLMo-1B and OLMo-7B, and expanding to local neighborhoods of up to 36K parameters provides no improvement. The failure is specific to Super Wei
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 48%Timing Trick Cuts Energy Used in LLM Training by Up to 14 Percent →
- PossiblePossibly related (embedding) · 47%Evaluate a model properly →
- PossiblePossibly related (embedding) · 46%Instead of decentralized training effort we should build the “One dataset” →
- PossiblePossibly related (embedding) · 46%H64LM: A 249M-parameter Mixture-of-Experts Transformer built from scratch in PyTorch [P] →
- PossiblePossibly related (embedding) · 46%Same GRPO recipe on three from-scratch LLMs (353M/316M/672M) gave three different outcomes, with no clean relationship to scale [P] →
- PossiblePossibly related (embedding) · 49%Continual Learning of Frontier Models for SovereignAI. Tech Report + Open Weights Model [R] →
- LinkedLinked via arxiv author · 85%Shreyas Subramanian →
“Super Weights in LLMs and the Failure of Selective Training”
