COSMI: COmpositional Synthesis of Multi-object Interactions
COSMI builds multi-object human interaction data by composing single-object captures into plausible combined actions.
The paper argues that most capture datasets miss everyday multi-object activity because recording combinations is too expensive. COSMI reuses contact-consistent single-object clips, mirrors and transfers them across bodies, then filters pairings with language-model and geometric checks. The resulting dataset has 222,000 sequences and 275 hours with up to five objects. A text-to-interaction diffusion transformer trained on it generalizes to unseen object and interaction combinations, outperforming baselines on text alignment and contact accuracy. HF Daily Papers' note
The paper argues that most capture datasets miss everyday multi-object activity because recording combinations is too expensive. COSMI reuses contact-consistent single-object clips, mirrors and transfers them across bodies, then filters pairings with language-model and geometric checks. The resulting dataset has 222,000 sequences and 275 hours with up to five objects. A text-to-interaction diffusion transformer trained on it generalizes to unseen object and interaction combinations, outperforming baselines on text alignment and contact accuracy. HF Daily Papers' note
score 4