Towards a Phonology-Informed Evaluation of Multilingual TTS
Neural TTS systems can sound natural across languages, but naturalness does not guarantee the preservation of sound contrasts that distinguish words from their grammatical forms. Standard metrics like MOS do not test for this. We propose a classifier-based framework that audits TTS output against language-specific phonological patterns using human speech as a benchmark. Testing Assamese advanced tongue root (ATR) vowel harmony with Meta's MMS TTS, we show that a classifier trained on human speech transfers to synthesized speech with minimal loss. The faithfulness audit reveals that [+ATR] mid
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via arxiv author · 85%Sneha Ray Barman →
“Towards a Phonology-Informed Evaluation of Multilingual TTS”
- LinkedLinked via arxiv author · 85%Neeraj Kumar Sharma →
“Towards a Phonology-Informed Evaluation of Multilingual TTS”
- LinkedLinked via arxiv author · 85%Shakuntala Mahanta →
“Towards a Phonology-Informed Evaluation of Multilingual TTS”
- PossiblePossibly related (embedding) · 48%Kyutai's Pocket TTS clones a voice from 5 seconds of audio, on CPU, under MIT. Benchmarked against Kokoro, Supertonic, and Inflect-Nano for Eng. TTS →
- FuzzyOverlapping authors or contributors · 62%ultralytics/yolov5 →
“Shared author/contributor keys: sharma”
- FuzzyOverlapping authors or contributors · 62%deepfakes/faceswap →
“Shared author/contributor keys: sharma”
