A Practical Investigation of Training-free Relaxed Speculative Decoding
Speculative decoding accelerates sampling from an autoregressive LLM by using a faster auxiliary model to draft tokens which are then verified in parallel by the LLM. Standard speculative decoding is lossless: its rejection and resampling steps exactly preserve the LLM's sampling distribution. Recent work argues that relaxing this strict guarantee can yield further speed-ups, controlled capability-speed trade-offs, or even capability gains. We practically investigate training-free relaxed speculative decoding techniques, unify existing approaches within a shared framework, benchmark them on co
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 67%lightseekorg/TorchSpec →
- PossiblePossibly related (embedding) · 63%sgl-project/SpecForge →
- PossiblePossibly related (embedding) · 59%[Research] JetSpec: Speculative Decoding with Parallel Tree Drafting Enables up to 9.64x Lossless LLM Inference Speedup with more than 1000TPS →
- PossiblePossibly related (embedding) · 59%DSpark: Speculative decoding accelerates LLM inference [pdf] →
- PossiblePossibly related (embedding) · 47%New Server Hopes to Break Through AI’s “Memory Wall” →
- LinkedLinked via arxiv author · 85%Guoxuan Xia →
“A Practical Investigation of Training-free Relaxed Speculative Decoding”
- LinkedLinked via arxiv author · 85%Luka Ribar →
“A Practical Investigation of Training-free Relaxed Speculative Decoding”
- LinkedLinked via arxiv author · 85%Paul Balanca →
“A Practical Investigation of Training-free Relaxed Speculative Decoding”
