WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS
While recent Large Language Model (LLM)-based Text-to-Speech (TTS) systems have achieved remarkable naturalness, they predominantly rely on implicit end-to-end generation paradigms, resulting in coarse-grained control. In scenarios demanding precise stylistic interventions and strict temporal alignment, such as audiobook narration and video dubbing, the inability to explicitly manipulate word-level acoustic attributes remains a critical bottleneck. This limitation is primarily amplified by the severe scarcity of fine-grained annotated datasets and the architectural challenge of integrating mul
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 54%lgy1027/matrix-live-diarizer →
- PossiblePossibly related (embedding) · 51%Text-to-Speech AI Edits Single Words Mid-Recording: ViiTorVoice Goes Open Source - Tech Times →
- PossiblePossibly related (embedding) · 51%Whisper-Lite →
- PossiblePossibly related (embedding) · 51%huggingface/speech-to-speech →
- PossiblePossibly related (embedding) · 50%Atomic-man007/Awesome_Multimodel_LLM →
- FuzzySimilar title/name (fuzzy) · 87%lllyasviel/ControlNet →
“Fuzzy title match (0.94): “WordVoice: Explicit and Decoupled Multi-Dimensional Word-Lev” ≈ “lllyasviel/ControlNet””
- FuzzySimilar title/name (fuzzy) · 87%lllyasviel/ControlNet-v1-1 →
“Fuzzy title match (0.94): “WordVoice: Explicit and Decoupled Multi-Dimensional Word-Lev” ≈ “lllyasviel/ControlNet-v1-1””
- FuzzyOverlapping authors or contributors · 62%Tongyi-MAI/Z-Image-Turbo →
“Shared author/contributor keys: mai”
