X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment
While large audio-language models have achieved remarkable progress in auditory perception, they still lag behind text-based large language models in deep logical reasoning, primarily due to the scarcity of high-quality audio reasoning data. To bridge this gap, we propose X$^3$-OPD, a cross-modal on-policy distillation framework that transfers reasoning capabilities from a powerful text teacher to an audio-language student. During training, the student generates reasoning trajectories conditioned on its own acoustic perception, while the teacher provides token-level guidance using matched text
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzyOverlapping authors or contributors · 62%HKUDS/LightRAG →
“Shared author/contributor keys: jin”
- FuzzyOverlapping authors or contributors · 62%keras-team/keras →
“Shared author/contributor keys: jin”
- FuzzyOverlapping authors or contributors · 62%DietrichGebert/ponytail →
“Shared author/contributor keys: cheng”
- LinkedLinked via arxiv author · 85%Dongjie Fu →
“X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment”
- LinkedLinked via arxiv author · 85%Di Cao →
“X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment”
- LinkedLinked via arxiv author · 85%Xize Cheng →
“X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment”
- LinkedLinked via arxiv author · 85%Zihan Zhang →
“X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment”
- LinkedLinked via arxiv author · 85%Wenxu Jia →
“X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment”
