From Scores to Evidence: Auditable Decisions Can Improve Speech Deepfake Detection
Speech deepfakes can mimic a speaker's voice convincingly enough to deceive listeners and automated systems. This has driven strong progress in speech deepfake detection, but most detectors still end with one score per utterance. That score is useful for ranking systems, yet it says little about why a borderline item should be trusted, deferred, or reviewed. Two utterances can fall in the same score band for different reasons, for example because passive and retrieval evidence disagree or because the keyed probe is unavailable. We ask whether the final decision can remain scalar without discar
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 49%Measuring benchmark optimization in speech recognition →
- PossiblePossibly related (embedding) · 48%Apple's new SpeechAnalyzer API, benchmarked against Whisper and its predecessor →
- PossiblePossibly related (embedding) · 48%Introducing Real World VoiceEQ: Measuring the human quality of voice AI →
- PossiblePossibly related (embedding) · 48%Is speech-to-text AI really reliable? - University of Cincinnati →
- PossiblePossibly related (embedding) · 47%The bottleneck for meeting transcription tools isn't accurate anymore, it's speaker attribution →
- FuzzySimilar title/name (fuzzy) · 87%huggingface/speech-to-speech →
“Fuzzy title match (0.94): “From Scores to Evidence: Auditable Decisions Can Improve Spe” ≈ “huggingface/speech-to-speech””
- LinkedLinked via arxiv author · 85%Mengzhe Geng →
“From Scores to Evidence: Auditable Decisions Can Improve Speech Deepfake Detection”
- LinkedLinked via arxiv author · 85%Yujia Lu →
“From Scores to Evidence: Auditable Decisions Can Improve Speech Deepfake Detection”
