Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning
Large Audio Language models (LALMs) have made rapid progress on acoustic understanding, yet they still struggle with fine-grained audio reasoning (e.g., recognizing event order, repetitions and duration). Existing post-training methods heavily rely on expensive external labels or provide only coarse semantic signals. To bridge this gap, we introduce Audio-Zero, the first label-free self-evolution framework in the field of LALMs that improves fine-grained auditory perception and reasoning. Audio-Zero constructs an auditory self-play game from unlabeled audio contrast pairs: most players hear a
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%modular/modular →
“Shared author/contributor keys: liu”
- FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow →
“Shared author/contributor keys: wang”
- LinkedLinked via arxiv author · 85%Siqian Tong →
“Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning”
- LinkedLinked via arxiv author · 85%Yaxuan Liu →
“Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning”
- LinkedLinked via arxiv author · 85%Chaozhuo Li →
“Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning”
- LinkedLinked via arxiv author · 85%Baolong Bi →
“Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning”
- LinkedLinked via arxiv author · 85%Yiwei Wang →
“Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning”
