BiSCo-LLM: Lookup-Free Binary Spherical Coding for Extreme Low-Bit Large Language Model Compression
Large language models (LLMs) are increasingly constrained by memory capacity, weight bandwidth, and checkpoint storage during deployment. Existing low-bit compression methods mainly follow two directions. Scalar or group-wise quantization is simple and compatible with efficient low-precision kernels, but its representation capacity becomes limited when the target budget approaches 2 bits per weight. Vector-quantized weight compression provides a richer block-level representation, but usually introduces explicit codebooks, index lookup, and additional storage accounting. This paper presents BiS
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 27%bitsandbytes-foundation/bitsandbytes →
“Possibly related via embedding similarity 0.59 (not asserted). Timestamp check: artifact slightly before paper (-2d).”
- PossiblePossibly related (embedding) · 49%Tencent/AngelSlim →
- PossiblePossibly related (embedding) · 48%TilelliLab/atome-lm →
- PossiblePossibly related (embedding) · 47%trvon/yams →
- PossiblePossibly related (embedding) · 47%New Server Hopes to Break Through AI’s “Memory Wall” →
- PossiblePossibly related (embedding) · 46%Show HN: misa77 - a codec that decodes 2x faster than LZ4 (at better ratios) →
- FuzzyOverlapping authors or contributors · 62%modular/modular →
“Shared author/contributor keys: liu”
- FuzzyOverlapping authors or contributors · 62%DietrichGebert/ponytail →
“Shared author/contributor keys: cheng”
