Finetuning Strategies for Querying Sounds by Vocal Imitation
This technical report describes our winning submission to the AES AIMLA 2025 Challenge on querying sound effects by vocal imitation. We investigate two complementary fine-tuning strategies: contrastive learning with a frozen, pretrained CED encoder, and joint contrastive-triplet learning with semi-hard negatives using a MobileNetV3 encoder. This report has been updated for posterity to include details released after the challenge.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 53%Introducing Real World VoiceEQ: Measuring the human quality of voice AI →
- PossiblePossibly related (embedding) · 52%Launch HN: Speko (YC S26) – OpenRouter for Voice AI →
- PossiblePossibly related (embedding) · 49%Kyutai's Pocket TTS clones a voice from 5 seconds of audio, on CPU, under MIT. Benchmarked against Kokoro, Supertonic, and Inflect-Nano for Eng. TTS →
- PossiblePossibly related (embedding) · 52%Best AI Voice Cloning in 2026: How to Clone Your Voice With AI →
- PossiblePossibly related (embedding) · 55%Granite Speech 5.0 Turbo CTC: Extremely Fast and Accurate Transcription →
- LinkedLinked via arxiv author · 85%Aditya Bhattacharjee →
“Finetuning Strategies for Querying Sounds by Vocal Imitation”
- LinkedLinked via arxiv author · 85%Christos Plachouras →
“Finetuning Strategies for Querying Sounds by Vocal Imitation”
- LinkedLinked via arxiv author · 85%Sungkyun Chang →
“Finetuning Strategies for Querying Sounds by Vocal Imitation”
