MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification
ARMDIL uses an MLLM to choose which vision model should handle each image.
The paper frames the problem as cross-dataset classification, where single-dataset specialists lose reliability across domains and difficulty levels. Its ensemble combines ResNets, self-supervised learners, and vision-language models trained over a unified label space. The router is competitive with specialized training-based routers, while prompt edits let it absorb new information more easily. The authors also point to natural-language routing traces as an interpretability gain. ArXiv · AI/CL/LG's note
The paper frames the problem as cross-dataset classification, where single-dataset specialists lose reliability across domains and difficulty levels. Its ensemble combines ResNets, self-supervised learners, and vision-language models trained over a unified label space. The router is competitive with specialized training-based routers, while prompt edits let it absorb new information more easily. The authors also point to natural-language routing traces as an interpretability gain. ArXiv · AI/CL/LG's note
score 4