newsHacker NewsTrust 72 · CommunityPublished 1mo agoLive · 1mo ago
DSpark: Speculative decoding accelerates LLM inference [pdf]
717points294comments
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownSpeculative decoding with draft models →
- LinkedLinked via unknownWhen are likely answers right? On Sequence Probability and Correctness in LLMs →
- LinkedLinked via unknownDepth Exploration for LLM Decoding →
- LinkedLinked via unknownBlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative Decoding →
- PossiblePossibly related (embedding) · 53%Kaden-Schutt/hipfire →
- PossiblePossibly related (embedding) · 52%sgl-project/SpecForge →
- PossiblePossibly related (embedding) · 55%alibaba/rtp-llm →
- PossiblePossibly related (embedding) · 56%dphnAI/aphrodite-engine →
Covers
Covers (incoming)
paperDepth Exploration for LLM DecodingpaperBlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative DecodingrepoKaden-Schutt/hipfirereposgl-project/SpecForgerepoalibaba/rtp-llmrepodphnAI/aphrodite-enginerepoguoqingbao/xinferrepolightseekorg/TorchSpecrepoEricLBuehler/mistral.rsrepodphnAI/sonarpaperDSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive GenerationpaperDominoTree: Conditional Tree-Structured Drafting with Domino for Speculative DecodingpaperA Practical Investigation of Training-free Relaxed Speculative DecodingrepoAarambhDevHub/aarambh-aipaperLess Experts, Faster Decoding: Cost-Aware Speculative Decoding for Mixture-of-Expertsrepofacebookresearch/LayerSkippaperAdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Draftersrepowarpfront/hipfire
Related across the graph
repofacebookresearch/LayerSkiprepolightseekorg/TorchSpecpaperWhen are likely answers right? On Sequence Probability and Correctness in LLMspaperBlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative DecodingpaperDominoTree: Conditional Tree-Structured Drafting with Domino for Speculative Decodingrepowarpfront/hipfirerepodphnAI/sonarrepoalibaba/rtp-llmreposgl-project/SpecForgepaperDSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive GenerationrepoAarambhDevHub/aarambh-aipaperDepth Exploration for LLM DecodingpaperAdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Draftersrepoguoqingbao/xinferrepoEricLBuehler/mistral.rsrepodphnAI/aphrodite-enginepaperA Practical Investigation of Training-free Relaxed Speculative DecodingrepoKaden-Schutt/hipfirepaperSpeculative decoding with draft modelspaperLess Experts, Faster Decoding: Cost-Aware Speculative Decoding for Mixture-of-Experts
