Read original ↗
paperarXivTrust 82 · PrimaryPublished 2d agoLive · 4m ago

MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification

Modern image classification models excel when trained on single task-specific datasets but often struggle to generalize across domains and difficulty levels. We propose ARMDIL, an Adaptive Router for Multi-Domain Image classification with LLMs. ARMDIL is an ensemble that uses a multimodal large language model (MLLM) agent to dynamically route each image to the most suitable vision backbone. Our diverse ensemble employs convolutional neural networks (ResNets), self-supervised representation learners (SSL), and vision-language models (VLMs), each trained on a unified label space constructed from

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzySimilar title/name (fuzzy) · 59%Tongyi-MAI/Z-Image-Turbo

    Fuzzy title match (0.73): “MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset” ≈ “Tongyi-MAI/Z-Image-Turbo”

  • LinkedLinked via arxiv author · 85%Daniel Perkins

    MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification

  • LinkedLinked via arxiv author · 85%John Squires

    MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification

  • LinkedLinked via arxiv author · 85%Janou Milligan

    MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification

  • LinkedLinked via arxiv author · 85%Chandra Raskoti

    MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification

  • LinkedLinked via arxiv author · 85%Linda Ungerboeck

    MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification

Has model

authored (incoming)

Related across the graph

Topics