Super-Tuning: From Activation-Aware Pruning to Sparse Fine-Tuning
Large language models (LLMs) remain expensive to fine-tune because full-parameter updates require substantial memory, compute, and per-task storage. We study whether saliency signals originally developed for pruning can be reused to choose where a model should adapt. We propose Super, a sparse parameter-efficient fine-tuning (PEFT) method that fixes a small trainable support using a Wanda-style activation-weighted magnitude score [Sun et al., 2023] computed from a calibration pass. We then introduce Supra, a hybrid adapter that combines this sparse update with LoRA while preserving a matched t
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 51%Breakthrough in long-context efficiency announced →
- PossiblePossibly related (embedding) · 50%attention-zoo →
- PossiblePossibly related (embedding) · 50%thu-pacman/chitu →
- PossiblePossibly related (embedding) · 48%chrisliu298/awesome-llm-unlearning →
- PossiblePossibly related (embedding) · 48%Looking for feedback on a small test SLM I built completely from scratch [P] →
- LinkedLinked via arxiv author · 85%Ivan Ilin →
“Super-Tuning: From Activation-Aware Pruning to Sparse Fine-Tuning”
- LinkedLinked via arxiv author · 85%Philip Zmushko →
“Super-Tuning: From Activation-Aware Pruning to Sparse Fine-Tuning”
- LinkedLinked via arxiv author · 85%Peter Richtárik →
“Super-Tuning: From Activation-Aware Pruning to Sparse Fine-Tuning”
