On-Policy Self-Distillation in Diffusion Models
Reinforcement learning can align diffusion models with human preferences and task-specific objectives, but endpoint rewards do not specify how an intermediate denoising prediction should change. We introduce DiffusionOPSD as an on-policy self-distillation framework that converts image-level reward guidance into explicit targets for clean-output predictions at sampled queries. At each outer iteration, a frozen behavior policy generates trajectories and supplies query states and anchors. Reward gradients construct bounded positive and negative targets around each anchor. The trainable policy fit
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 51%AI models get convenient amnesia about source material as they grow, MIT boffins find →
- FuzzySimilar title/name (fuzzy) · 59%CompVis/stable-diffusion-v1-4 →
“Fuzzy title match (0.73): “On-Policy Self-Distillation in Diffusion Models” ≈ “CompVis/stable-diffusion-v1-4””
- FuzzyOverlapping authors or contributors · 62%Kong/kong →
“Shared author/contributor keys: kong”
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%google-research/google-research →
“Shared author/contributor keys: sun”
- FuzzyOverlapping authors or contributors · 62%modular/modular →
“Shared author/contributor keys: liu”
- FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow →
“Shared author/contributor keys: wang”
- LinkedLinked via arxiv author · 85%Xiaowei Zhou →
“On-Policy Self-Distillation in Diffusion Models”
