Megadose Built for builders and researchers.

Multimodal Flow: Unified Flow Modeling of Language and Vision in Embedding Spaces

· HF Daily Papers ·
The paper proposes a single continuous flow model for both text and images, avoiding discrete visual tokenization and split generation objectives.

Multimodal Flow represents text blocks and images as ordered continuous “hyperchunks,” then trains one chunk-causal flow backbone with Flow Matching. The authors say joint attention handles cross-modal interaction, while modality-specific feed-forward layers handle each modality’s processing. Their MF-1 runs at 0.6B, 1.2B, and 1.6B scales, with continued pretraining improving multimodal results. With 150B pretraining tokens, they report competitive scores on GenEval, DPG-Bench, VQAv2, MMBench, and POPE, and say it beats representative hybrid and discrete models under matched budgets. HF Daily Papers' note

score 5

Categories: Research