newsReddit r/LocalLLaMATrust 58 · CommunityPublished 1mo agoLive · 1mo ago
[Research] JetSpec: Speculative Decoding with Parallel Tree Drafting Enables up to 9.64x Lossless LLM Inference Speedup with more than 1000TPS
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownSpeculative decoding with draft models →
- LinkedLinked via unknownDepth Exploration for LLM Decoding →
- LinkedLinked via unknownBlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative Decoding →
- PossiblePossibly related (embedding) · 53%sgl-project/SpecForge →
- PossiblePossibly related (embedding) · 59%lightseekorg/TorchSpec →
- PossiblePossibly related (embedding) · 59%DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation →
- PossiblePossibly related (embedding) · 55%DominoTree: Conditional Tree-Structured Drafting with Domino for Speculative Decoding →
- PossiblePossibly related (embedding) · 59%A Practical Investigation of Training-free Relaxed Speculative Decoding →
Covers
Covers (incoming)
paperDepth Exploration for LLM DecodingpaperBlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative Decodingreposgl-project/SpecForgerepolightseekorg/TorchSpecpaperDSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive GenerationpaperDominoTree: Conditional Tree-Structured Drafting with Domino for Speculative DecodingpaperA Practical Investigation of Training-free Relaxed Speculative DecodingpaperLess Experts, Faster Decoding: Cost-Aware Speculative Decoding for Mixture-of-ExpertspaperWindowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token ContextpaperDARTree: Speculative Diffusion Decoding with Autoregressive Draft Trees
Related across the graph
paperDARTree: Speculative Diffusion Decoding with Autoregressive Draft Treesrepolightseekorg/TorchSpecpaperBlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative DecodingpaperDominoTree: Conditional Tree-Structured Drafting with Domino for Speculative Decodingreposgl-project/SpecForgepaperDSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive GenerationpaperDepth Exploration for LLM DecodingpaperA Practical Investigation of Training-free Relaxed Speculative DecodingpaperSpeculative decoding with draft modelspaperLess Experts, Faster Decoding: Cost-Aware Speculative Decoding for Mixture-of-ExpertspaperWindowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token Context
