paperarXivTrust 82 · PrimaryPublished 1mo agoLive · 3mo ago
Speculative decoding with draft models
Accelerating generation by drafting tokens with a small model.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownToken →
- LinkedLinked via unknownDSpark: Speculative decoding accelerates LLM inference [pdf] →
- LinkedLinked via unknownNew Server Hopes to Break Through AI’s “Memory Wall” →
- PossiblePossibly related (embedding) · 63%sgl-project/SpecForge →
- PossiblePossibly related (embedding) · 66%lightseekorg/TorchSpec →
- PossiblePossibly related (embedding) · 50%Faster LLMs Inference: Speculative Decoding Explained - YouTube →
Related to (incoming)
Covers (incoming)
news[Research] JetSpec: Speculative Decoding with Parallel Tree Drafting Enables up to 9.64x Lossless LLM Inference Speedup with more than 1000TPSnewsDSpark: Speculative decoding accelerates LLM inference [pdf]newsNew Server Hopes to Break Through AI’s “Memory Wall”newsFaster LLMs Inference: Speculative Decoding Explained - YouTube
Implements (incoming)
Related across the graph
repolightseekorg/TorchSpecnewsNew Server Hopes to Break Through AI’s “Memory Wall”news[Research] JetSpec: Speculative Decoding with Parallel Tree Drafting Enables up to 9.64x Lossless LLM Inference Speedup with more than 1000TPSreposgl-project/SpecForgenewsDSpark: Speculative decoding accelerates LLM inference [pdf]glossary_termTokennewsFaster LLMs Inference: Speculative Decoding Explained - YouTube
