huggingface/speech-to-speech
Build local voice agents with open-source models
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 61%Looking for open-source AI meeting note-takers (like Fathom, Fireflies, Notion AI) →
- PossiblePossibly related (embedding) · 54%How Loka Built a Natural, Low-Latency Voice Agent with Amazon Nova 2 Sonic →
- FuzzySimilar title/name (fuzzy) · 87%Auditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech Evaluation →
“Fuzzy title match (0.94): “Auditing Protocol-Level Shortcuts in Large Audio Language Mo” ≈ “huggingface/speech-to-speech””
- FuzzySimilar title/name (fuzzy) · 87%Self-supervised Speech Comparison for L2 Phone, Rhythm, and Intonation Scoring →
“Fuzzy title match (0.94): “Self-supervised Speech Comparison for L2 Phone, Rhythm, and ” ≈ “huggingface/speech-to-speech””
- FuzzySimilar title/name (fuzzy) · 87%BlueMagpie-TTS: A Token-Efficient Tokenizer, Language Model, and TTS for Taiwanese-Accent Code-Switching Speech →
“Fuzzy title match (0.94): “BlueMagpie-TTS: A Token-Efficient Tokenizer, Language Model,” ≈ “huggingface/speech-to-speech””
- FuzzySimilar title/name (fuzzy) · 84%SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models →
“Fuzzy title match (0.92): “SPEARBench: A Benchmark for Naturalness Evaluation in Stream” ≈ “huggingface/speech-to-speech””
- FuzzySimilar title/name (fuzzy) · 87%From Black-Box to Clinical Insight: A Multi-Stage Explainable Framework for Speech-Based Cognitive Impairment Detection →
“Fuzzy title match (0.94): “From Black-Box to Clinical Insight: A Multi-Stage Explainabl” ≈ “huggingface/speech-to-speech””
- FuzzySimilar title/name (fuzzy) · 87%Comparing Human and Automatic Recognition of Dutch Dysarthric Continuous Speech: A Case Study →
“Fuzzy title match (0.94): “Comparing Human and Automatic Recognition of Dutch Dysarthri” ≈ “huggingface/speech-to-speech””
