Read original ↗
paperarXivTrust 82 · PrimaryPublished 28d agoLive · 26d ago

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances

Discrete speech tokenizers aim to disentangle semantic from acoustic information, yet targets from self-supervised learning (SSL) models like HuBERT retain non-linguistic variation: speaker identity, prosody, and channel conditions leak into the tokens, inflating entropy. Our key insight is that when enough speakers utter the same words under varying conditions, linguistic content is the only shared factor. We propose PINT (Parallel INvariant Tokenization), which fine-tunes an SSL encoder with alignment losses across parallel utterances and augmentations to distill this shared residual. PINT c

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzySimilar title/name (fuzzy) · 87%huggingface/speech-to-speech

    Fuzzy title match (0.94): “Content is What Remains: Invariant Speech Tokenization from ” ≈ “huggingface/speech-to-speech”

  • LinkedLinked via arxiv author · 85%Laurin Wagner

    Content is What Remains: Invariant Speech Tokenization from Parallel Utterances

  • LinkedLinked via arxiv author · 85%Bernhard Thallinger

    Content is What Remains: Invariant Speech Tokenization from Parallel Utterances

  • LinkedLinked via arxiv author · 85%Miroslav Stankovic

    Content is What Remains: Invariant Speech Tokenization from Parallel Utterances

  • LinkedLinked via arxiv author · 85%Mario Zusag

    Content is What Remains: Invariant Speech Tokenization from Parallel Utterances

Implements (incoming)

authored (incoming)

Related across the graph

Topics