Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs
Serving large language models cheaply increasingly means shipping models that are both structurally compressed to a fraction of their parameters and quantized to 4 bits. Together these steps degrade reasoning, mathematics, coding, and long-context behavior enough to require a recovery, or healing, stage before deployment. The default recipe, quantization-aware training (QAT), re-fits the compressed, quantized model to hard labels; in our pipeline it converged slowly and collapsed past its peak. We adopted Quantization-Aware Healing (QAH) instead. Because a structurally compressed model is neve
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 57%I developed my own quantized LLM from scratch, trained on 30B tokens, deploys in 60 MB [R] →
- PossiblePossibly related (embedding) · 58%Quantization →
- PossiblePossibly related (embedding) · 55%[Paper] Statistically-Lossless Quantization of Large Language Models →
- PossiblePossibly related (embedding) · 54%[R] Statistically-Lossless Quantization of Large Language Models →
- FuzzyOverlapping authors or contributors · 62%stefan-jansen/machine-learning-for-trading →
“Shared author/contributor keys: jansen”
- PossiblePossibly related (embedding) · 70%Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original →
- LinkedLinked via arxiv author · 85%Bakbergen Ryskulov →
“Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs”
- LinkedLinked via arxiv author · 85%Iker García-Ferrero →
“Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs”
