MoTE: Mixture of Task Experts for Multi-Task Video Understanding
MoTE routes each video example through a task-specific decoder expert while keeping the shared multimodal backbone.
The paper targets procedural video-language tasks that draw on the same visual evidence but need different behaviors, such as recognition, forecasting, and procedure prediction. Its VideoLLM-MoTE setup uses explicit task routes across five COIN benchmarks, activating about 2B LLM parameters per sample. The authors report higher average top-1 accuracy than recent VideoLLM baselines, and better results than dense all-expert activation or learned sparse-routing controls under the same topology. HF Daily Papers' note
The paper targets procedural video-language tasks that draw on the same visual evidence but need different behaviors, such as recognition, forecasting, and procedure prediction. Its VideoLLM-MoTE setup uses explicit task routes across five COIN benchmarks, activating about 2B LLM parameters per sample. The authors report higher average top-1 accuracy than recent VideoLLM baselines, and better results than dense all-expert activation or learned sparse-routing controls under the same topology. HF Daily Papers' note
score 4