Understanding Evaluation Illusion in Diffusion Large Language Models
Despite the capability of parallel decoding, diffusion large language models (dLLMs) require many denoising steps to maintain generation quality, motivating recent research on efficient decoding strategies. However, existing studies have reported inconsistent evaluation results even under seemingly identical evaluation settings, risking biased conclusions about dLLM decoding methods. To understand this evaluation concern, we conduct a rigorous evaluation of current decoding methods for dLLMs across diverse evaluation settings. Surprisingly, our analysis reveals that the ranking of decoding met
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownminimal-diffusion-lm →
- LinkedLinked via unknownDiffusionGemma: 4x faster text generation →
- LinkedLinked via unknownTransformer →
- FuzzySimilar title/name (fuzzy) · 59%CompVis/stable-diffusion-v1-4 →
“Fuzzy title match (0.73): “Understanding Evaluation Illusion in Diffusion Large Languag” ≈ “CompVis/stable-diffusion-v1-4””
- FuzzySimilar title/name (fuzzy) · 59%stabilityai/stable-diffusion-3.5-large →
“Fuzzy title match (0.73): “Understanding Evaluation Illusion in Diffusion Large Languag” ≈ “stabilityai/stable-diffusion-3.5-large””
- PossiblePossibly related (embedding) · 47%Contrastive Decoding Diffing (CDD): recovering verbatim finetuning data from logits alone, no weight access needed[R] →
- PossiblePossibly related (embedding) · 51%Scaling Properties of Continuous Diffusion Spoken Language Models - Apple Machine Learning Research →
