Phoneme- and Word-Level Metrics Using Self-Supervised Speech Representations for Forced Alignment Evaluation
Forced alignment evaluation typically requires manually annotated timestamps, limiting large-scale and multilingual analysis. We introduce two corpus-level metrics based on self-supervised (SSL) speech representations for reference-free forced alignment evaluation: Phoneme-Cluster Mutual Information (PCMI) and Word Acoustic Consistency Score (WACS). PCMI measures agreement between aligned phoneme labels and clusters induced from SSL-speech representations, while WACS measures consistency of repeated word realizations using dynamic time warping similarity between word representation sequences.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 49%Measuring benchmark optimization in speech recognition →
- PossiblePossibly related (embedding) · 45%Granite Speech 5.0 Turbo CTC: Extremely Fast and Accurate Transcription →
- FuzzySimilar title/name (fuzzy) · 87%huggingface/speech-to-speech →
“Fuzzy title match (0.94): “Phoneme- and Word-Level Metrics Using Self-Supervised Speech” ≈ “huggingface/speech-to-speech””
- LinkedLinked via arxiv author · 85%V. S. D. S. Mahesh Akavarapu →
“Phoneme- and Word-Level Metrics Using Self-Supervised Speech Representations for Forced Alignment Evaluation”
- LinkedLinked via arxiv author · 85%Michael Daniel →
“Phoneme- and Word-Level Metrics Using Self-Supervised Speech Representations for Forced Alignment Evaluation”
- LinkedLinked via arxiv author · 85%Gerhard Jäger →
“Phoneme- and Word-Level Metrics Using Self-Supervised Speech Representations for Forced Alignment Evaluation”
