Induction in Both Directions: A Mechanistic Analysis of In-Context Learning in Masked Diffusion Language Models
While the internal mechanisms of autoregressive (AR) transformers have been studied extensively, much less is known about diffusion language models (DLMs), an emerging alternative that generates text by iterative denoising. In this work, we study how DLMs implement induction, a mechanism behind in-context learning in which the model finds a repeated context and copies the token that followed it. Our analysis compares attention-only AR models and absorbing-mask DLMs with matched architectures. We find that DLMs learn a bidirectional induction circuit, where previous-token and next-token heads w
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 59%stabilityai/stable-diffusion-xl-base-1.0 →
“Fuzzy title match (0.73): “Induction in Both Directions: A Mechanistic Analysis of In-C” ≈ “stabilityai/stable-diffusion-xl-base-1.0””
- FuzzySimilar title/name (fuzzy) · 59%CompVis/stable-diffusion-v1-4 →
“Fuzzy title match (0.73): “Induction in Both Directions: A Mechanistic Analysis of In-C” ≈ “CompVis/stable-diffusion-v1-4””
- FuzzySimilar title/name (fuzzy) · 59%stabilityai/stable-diffusion-3.5-large →
“Fuzzy title match (0.73): “Induction in Both Directions: A Mechanistic Analysis of In-C” ≈ “stabilityai/stable-diffusion-3.5-large””
- PossiblePossibly related (embedding) · 61%What if context compression is a diffusion noise function? Proposal + honest results from untrained-model experiments [R] →
- PossiblePossibly related (embedding) · 57%Transformer →
- PossiblePossibly related (embedding) · 57%Learning Unmasking Policies for Diffusion Language Models - Apple Machine Learning Research →
- PossiblePossibly related (embedding) · 51%DiffusionGemma: 4x faster text generation →
- FuzzySimilar title/name (fuzzy) · 59%aymericdamien/TopDeepLearning →
“Fuzzy title match (0.73): “Induction in Both Directions: A Mechanistic Analysis of In-C” ≈ “aymericdamien/TopDeepLearning””
