Read original ↗
paperarXivTrust 82 · PrimaryPublished 8d agoLive · 7d ago

IDeaL: Data-Free Multi-Teacher Distillation via Improved Dead Leaves

Multi-teacher distillation has emerged as a way to combine complementary teacher models into a single student model that exhibits the strengths of all its teachers. The student is trained to mimic the output of the teachers on a set of images, typically the union of the individual teacher's training sets, assuming this data is available. In this paper, we question that assumption and explore alternative options. We first study how far one can go when distilling from teachers fed with different types of noise. Then, we show that information contained in the teachers can be leveraged to tailor t

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • LinkedLinked via arxiv author · 85%Feyza Yavuz

    IDeaL: Data-Free Multi-Teacher Distillation via Improved Dead Leaves

  • LinkedLinked via arxiv author · 85%Mert Bülent Sarıyıldız

    IDeaL: Data-Free Multi-Teacher Distillation via Improved Dead Leaves

  • LinkedLinked via arxiv author · 85%Diane Larlus

    IDeaL: Data-Free Multi-Teacher Distillation via Improved Dead Leaves

authored (incoming)

Related across the graph

Topics