Streaming Neural Speech Codecs through Time-Invariant Representations
Neural speech codecs are increasingly used as intermediate representations in codec-based speech generation systems. TiCodec introduces a factorized representation that separates time-varying speech content from time-invariant information through a Time-Invariant Representation Extraction (TIRE) module, potentially reducing the amount of information that must be modeled at the frame-level. In this work, we investigate the nature of the information captured by TIRE representations and their suitability for low-latency speech processing. Using a series of probing tasks, we analyze the influenc
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 51%lgy1027/matrix-live-diarizer →
- LinkedLinked via arxiv author · 85%Kélian Estève →
“Streaming Neural Speech Codecs through Time-Invariant Representations”
- LinkedLinked via arxiv author · 85%Salima Mhdaffar →
“Streaming Neural Speech Codecs through Time-Invariant Representations”
- LinkedLinked via arxiv author · 85%Mickael Rouvier →
“Streaming Neural Speech Codecs through Time-Invariant Representations”
- LinkedLinked via arxiv author · 85%Richard Dufour →
“Streaming Neural Speech Codecs through Time-Invariant Representations”
- LinkedLinked via arxiv author · 85%Yannick Estève →
“Streaming Neural Speech Codecs through Time-Invariant Representations”
- PossiblePossibly related (embedding) · 46%[audio.cpp] 10 hours of audio generated in 3 minutes on RTX 5090 (demo included)! C++/GGML based Supertonic 3, MOSS-TTS, IndexTTS2, and Irodori-TTS released →
- FuzzySimilar title/name (fuzzy) · 87%huggingface/speech-to-speech →
“Fuzzy title match (0.94): “Streaming Neural Speech Codecs through Time-Invariant Repres” ≈ “huggingface/speech-to-speech””
