AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters
Speculative decoding, in which a lightweight draft model first generates a draft sequence that is then verified in parallel by the target model, has become a prevalent paradigm for accelerating large language model inference. Recent work such as DFlash further boosts drafting efficiency by leveraging diffusion drafters, whose parallel denoising mechanism enables draft generation in a single forward pass. In this work, we uncover a central pitfall of diffusion drafters: bidirectional attention is a double-edged sword. On one hand, it endows the model with parallel generation and global contextu
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 53%DSpark: Speculative decoding accelerates LLM inference [pdf] →
- PossiblePossibly related (embedding) · 53%What if context compression is a diffusion noise function? Proposal + honest results from untrained-model experiments [R] →
- LinkedLinked via arxiv author · 85%Yu-Yang Qian →
“AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters”
- LinkedLinked via arxiv author · 85%Hao-Cong Wu →
“AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters”
- LinkedLinked via arxiv author · 85%Peter Yichen Chen →
“AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters”
- LinkedLinked via arxiv author · 85%Jiacheng Sun →
“AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters”
- LinkedLinked via arxiv author · 85%Zhenhua Dong →
“AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters”
- LinkedLinked via arxiv author · 85%Peng Zhao →
“AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters”
