Audio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Model
Automatic speech recognition is dominated by autoregressive decoders that emit one token at a time. We ask whether a discrete diffusion language model can transcribe speech instead, refining a whole transcript in parallel over a small number of denoising steps. We train an audio-native interface for DiffusionGemma, a 26B mixture-of-experts model that generates text by uniform, random-token discrete diffusion rather than the absorbing-mask scheme common to recent diffusion language models. A frozen Whisper encoder supplies acoustic features, a lightweight projector maps them into the model embe
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 58%Scaling Properties of Continuous Diffusion Spoken Language Models - Apple Machine Learning Research →
- PossiblePossibly related (embedding) · 54%minimal-diffusion-lm →
- PossiblePossibly related (embedding) · 52%Whisper-Lite →
- PossiblePossibly related (embedding) · 51%jaswon/osu-dreamer →
- PossiblePossibly related (embedding) · 50%Transformer →
- FuzzySimilar title/name (fuzzy) · 59%stabilityai/stable-diffusion-xl-base-1.0 →
“Fuzzy title match (0.73): “Audio-Native Speech Recognition with a Frozen Discrete-Diffu” ≈ “stabilityai/stable-diffusion-xl-base-1.0””
- FuzzySimilar title/name (fuzzy) · 59%CompVis/stable-diffusion-v1-4 →
“Fuzzy title match (0.73): “Audio-Native Speech Recognition with a Frozen Discrete-Diffu” ≈ “CompVis/stable-diffusion-v1-4””
- FuzzySimilar title/name (fuzzy) · 59%stabilityai/stable-diffusion-3.5-large →
“Fuzzy title match (0.73): “Audio-Native Speech Recognition with a Frozen Discrete-Diffu” ≈ “stabilityai/stable-diffusion-3.5-large””
