Megadose AI progress, ranked and analyzed.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities

· HF Daily Papers ·
Modus tests whether one decoder-only model can handle arbitrary inputs and outputs across modalities.

The paper frames any-to-any modeling as a single-network setup where every modality can serve as both input and target. Its claim is that Modus avoids modality-specific heads, losses, and task pipelines while using the strengths of decoder-only pretraining. The authors say this enables chained generation through intermediate modalities and self-checking by scoring outputs against another generated modality. They report competitive results against specialist and multitask baselines across benchmarks, with materials open-sourced. HF Daily Papers' note

score 5

Categories: Research