IDeaL: Data-Free Multi-Teacher Distillation via Improved Dead Leaves
Multi-teacher distillation has emerged as a way to combine complementary teacher models into a single student model that exhibits the strengths of all its teachers. The student is trained to mimic the output of the teachers on a set of images, typically the union of the individual teacher's training sets, assuming this data is available. In this paper, we question that assumption and explore alternative options. We first study how far one can go when distilling from teachers fed with different types of noise. Then, we show that information contained in the teachers can be leveraged to tailor t
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via arxiv author · 85%Feyza Yavuz →
“IDeaL: Data-Free Multi-Teacher Distillation via Improved Dead Leaves”
- LinkedLinked via arxiv author · 85%Mert Bülent Sarıyıldız →
“IDeaL: Data-Free Multi-Teacher Distillation via Improved Dead Leaves”
- LinkedLinked via arxiv author · 85%Diane Larlus →
“IDeaL: Data-Free Multi-Teacher Distillation via Improved Dead Leaves”
