Read original ↗
paperarXivTrust 82 · PrimaryPublished 6d agoLive · 4d ago

Trace-Based On-Policy Distillation for Masked Diffusion Language Models

Diffusion large language models (dLLMs) are a promising alternative to autoregressive generation. However, reasoning-oriented post-training for dLLMs remains challenging. Supervised fine-tuning (SFT) for dLLMs requires dense but often off-policy masked states, while reinforcement learning (RL) relies on sparse rewards or value modeling. This paper proposes \textbf{trace-based on-policy distillation (TOPD)}, a teacher-supervised framework that transfers reasoning ability to a target dLLM without reward estimation. The key idea is to supervise a dLLM on its own denoising trajectory, focusing on

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzySimilar title/name (fuzzy) · 59%CompVis/stable-diffusion-v1-4

    Fuzzy title match (0.73): “Trace-Based On-Policy Distillation for Masked Diffusion Lang” ≈ “CompVis/stable-diffusion-v1-4”

  • FuzzySimilar title/name (fuzzy) · 59%stabilityai/stable-diffusion-3.5-large

    Fuzzy title match (0.73): “Trace-Based On-Policy Distillation for Masked Diffusion Lang” ≈ “stabilityai/stable-diffusion-3.5-large”

  • FuzzyOverlapping authors or contributors · 62%modular/modular

    Shared author/contributor keys: liu

  • LinkedLinked via arxiv author · 85%Haolin Ren

    Trace-Based On-Policy Distillation for Masked Diffusion Language Models

  • LinkedLinked via arxiv author · 85%Ziyang Huang

    Trace-Based On-Policy Distillation for Masked Diffusion Language Models

  • LinkedLinked via arxiv author · 85%Chenhao Yuan

    Trace-Based On-Policy Distillation for Masked Diffusion Language Models

  • LinkedLinked via arxiv author · 85%Minjun Zhao

    Trace-Based On-Policy Distillation for Masked Diffusion Language Models

  • LinkedLinked via arxiv author · 85%Kang Liu

    Trace-Based On-Policy Distillation for Masked Diffusion Language Models

Has model

Implements (incoming)

authored (incoming)

Related across the graph

Topics