The Illusion of Equivalency: Statistical Characterization of Quantization Effects in LLMs
Post-training quantization is widely used to deploy large language models in resource-constrained settings, yet its evaluation relies almost exclusively on accuracy and perplexity. We show that these metrics fail to capture behavioral changes induced by quantization. We introduce correctness agreement, a decision-level metric that measures overlap in correct predictions between a base model and its quantized variants, independent of absolute accuracy. Across multiple models and quantization schemes from 8-bit to 2-bit, we find that behavioral divergence emerges under moderate quantization even
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 54%bitsandbytes-foundation/bitsandbytes →
- PossiblePossibly related (embedding) · 52%Quantization →
- PossiblePossibly related (embedding) · 51%chrisliu298/awesome-llm-unlearning →
- PossiblePossibly related (embedding) · 49%thu-pacman/chitu →
- PossiblePossibly related (embedding) · 47%cimeister/tokenizer-intrinsic-evals →
- PossiblePossibly related (embedding) · 46%LLMs Are Not (Consistently) Bayesian: Quantifying Internal (In)consistencies of LLMs’ Probabilistic Beliefs - Apple Machine Learning Research →
- LinkedLinked via arxiv author · 85%Baha Rababah →
“The Illusion of Equivalency: Statistical Characterization of Quantization Effects in LLMs”
- LinkedLinked via arxiv author · 85%Cuneyt Gurcan Akcora →
“The Illusion of Equivalency: Statistical Characterization of Quantization Effects in LLMs”
