OUTLETS: Output-Length Prediction from Speculative Decoding Backbones
The heavy-tailed distribution of output lengths in Large Language Model (LLM) serving poses major challenges for resource provisioning and cluster scheduling. Although output-length prediction can mitigate these issues, existing approaches have key drawbacks: external proxy models add substantial latency and often have limited fidelity, whereas internal state-based methods are efficient but rely on shallow probes of current model states. We identify a structural connection between speculative decoding (SD) and length prediction: latent representations produced by the draft decoder in advanced
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 52%Free LM Studio Accelerates LLM Inference with Three Speculative Decoding Methods - finance.biggo.com →
- PossiblePossibly related (embedding) · 50%[Research] JetSpec: Speculative Decoding with Parallel Tree Drafting Enables up to 9.64x Lossless LLM Inference Speedup with more than 1000TPS →
- FuzzyOverlapping authors or contributors · 62%modular/modular →
“Shared author/contributor keys: liu”
- FuzzyOverlapping authors or contributors · 62%sgl-project/sglang →
“Shared author/contributor keys: zhou”
- FuzzyOverlapping authors or contributors · 62%google-research/google-research →
“Shared author/contributor keys: sun”
- LinkedLinked via arxiv author · 85%Weihuang Wen →
“OUTLETS: Output-Length Prediction from Speculative Decoding Backbones”
- LinkedLinked via arxiv author · 85%Yingying Liu →
“OUTLETS: Output-Length Prediction from Speculative Decoding Backbones”
- LinkedLinked via arxiv author · 85%Yichuan Liu →
“OUTLETS: Output-Length Prediction from Speculative Decoding Backbones”
