Tokenizer Transplantation: Mitigating Autoregressive Collapse in Edge-Efficient Bengali ASR
Lightweight speech recognition models are critical for edge deployment, yet highly optimized architectures like Moonshine often fail on morphologically rich, non-Latin languages such as Bengali. This study identifies the root cause of this failure as the model's English-centric byte-level tokenizer, which fragments Bengali words into high-fertility byte chains and triggers catastrophic autoregressive collapse during inference. To resolve this, a novel vocabulary transplantation pipeline is proposed to replace the decoder vocabulary with the native-script BanglaBERT WordPiece vocabulary and res
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 48%Transformer →
- PossiblePossibly related (embedding) · 47%thu-pacman/chitu →
- PossiblePossibly related (embedding) · 46%bitsandbytes-foundation/bitsandbytes →
- LinkedLinked via arxiv author · 85%Sanjid Hasan →
“Tokenizer Transplantation: Mitigating Autoregressive Collapse in Edge-Efficient Bengali ASR”
- LinkedLinked via arxiv author · 85%Md. Abdur Rahman →
“Tokenizer Transplantation: Mitigating Autoregressive Collapse in Edge-Efficient Bengali ASR”
