Leading-Silence Augmentation and Multi-Stage Synthetic Supervision for the Second MLC-SLM Challenge
The second Multilingual Conversational Speech Language Model (MLC-SLM) Challenge evaluates two tasks over complete, unsegmented multilingual conversations: speaker diarization and recognition (Task 1) and conversational speech understanding (Task 2). Neither task provides oracle utterance boundaries or speaker labels at evaluation, and Task 2 provides no question-answer training set. For Task 1, we fine-tune VibeVoice-ASR-7B with random leading-silence cropping, consistent timestamp correction, and an exponential moving average (EMA) training strategy. For Task 2, we construct synthetic questi
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 51%Looking for feedback on a small test SLM I built completely from scratch [P] →
- PossiblePossibly related (embedding) · 50%Introducing Real World VoiceEQ: Measuring the human quality of voice AI →
- FuzzySimilar title/name (fuzzy) · 84%roboflow/supervision →
“Fuzzy title match (0.92): “Leading-Silence Augmentation and Multi-Stage Synthetic Super” ≈ “roboflow/supervision””
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%sgl-project/sglang →
“Shared author/contributor keys: zhou”
- FuzzyOverlapping authors or contributors · 62%google-research/google-research →
“Shared author/contributor keys: sun”
- FuzzyOverlapping authors or contributors · 62%modular/modular →
“Shared author/contributor keys: liu”
- LinkedLinked via arxiv author · 85%Kexin Shi →
“Leading-Silence Augmentation and Multi-Stage Synthetic Supervision for the Second MLC-SLM Challenge”
