DominoTree: Conditional Tree-Structured Drafting with Domino for Speculative Decoding
Speculative decoding accelerates LLM inference by drafting several tokens and verifying them in parallel. Block-diffusion drafters such as DFlash produce a draft block in one pass but model only per-position marginals; best-first tree methods such as DDTree expand candidate trees from those marginals. The released Domino drafter adds a GRU-based causal correction that makes each draft token's distribution path-dependent, a structure DDTree's factorized formulation cannot represent. We introduce DominoTree, a training-free best-first draft tree scored by Domino's conditional, non-factoriz
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 57%lightseekorg/TorchSpec →
- PossiblePossibly related (embedding) · 55%[Research] JetSpec: Speculative Decoding with Parallel Tree Drafting Enables up to 9.64x Lossless LLM Inference Speedup with more than 1000TPS →
- PossiblePossibly related (embedding) · 54%sgl-project/SpecForge →
- PossiblePossibly related (embedding) · 54%DSpark: Speculative decoding accelerates LLM inference [pdf] →
- PossiblePossibly related (embedding) · 48%EricLBuehler/mistral.rs →
- LinkedLinked via arxiv author · 85%Saw S. Lin →
“DominoTree: Conditional Tree-Structured Drafting with Domino for Speculative Decoding”
- LinkedLinked via arxiv author · 85%Jyh-Shing Roger Jang →
“DominoTree: Conditional Tree-Structured Drafting with Domino for Speculative Decoding”
- FuzzyOverlapping authors or contributors · 62%hiyouga/LlamaFactory →
“Shared author/contributor keys: lin”
