Content is What Remains: Invariant Speech Tokenization from Parallel Utterances
Discrete speech tokenizers aim to disentangle semantic from acoustic information, yet targets from self-supervised learning (SSL) models like HuBERT retain non-linguistic variation: speaker identity, prosody, and channel conditions leak into the tokens, inflating entropy. Our key insight is that when enough speakers utter the same words under varying conditions, linguistic content is the only shared factor. We propose PINT (Parallel INvariant Tokenization), which fine-tunes an SSL encoder with alignment losses across parallel utterances and augmentations to distill this shared residual. PINT c
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 87%huggingface/speech-to-speech →
“Fuzzy title match (0.94): “Content is What Remains: Invariant Speech Tokenization from ” ≈ “huggingface/speech-to-speech””
- LinkedLinked via arxiv author · 85%Laurin Wagner →
“Content is What Remains: Invariant Speech Tokenization from Parallel Utterances”
- LinkedLinked via arxiv author · 85%Bernhard Thallinger →
“Content is What Remains: Invariant Speech Tokenization from Parallel Utterances”
- LinkedLinked via arxiv author · 85%Miroslav Stankovic →
“Content is What Remains: Invariant Speech Tokenization from Parallel Utterances”
- LinkedLinked via arxiv author · 85%Mario Zusag →
“Content is What Remains: Invariant Speech Tokenization from Parallel Utterances”
