Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
Mage-VL cuts visual tokens by more than 75% by reading video through codec signals instead of uniform frame sampling.
The paper says its Mage-ViT tokenizer selects dynamic, high-entropy regions using motion vectors and residual energy across I- and P-frames. The model is built for real-time multimodal understanding, with a lightweight event gate feeding a causal decoder for streaming perception. In evaluations, Mage-VL-4B matches Qwen3-VL-4B on static tasks, improves on video and spatial reasoning, and reaches up to 3.5x faster wall-clock inference. HF Daily Papers' note
The paper says its Mage-ViT tokenizer selects dynamic, high-entropy regions using motion vectors and residual energy across I- and P-frames. The model is built for real-time multimodal understanding, with a lightweight event gate feeding a causal decoder for streaming perception. In evaluations, Mage-VL-4B matches Qwen3-VL-4B on static tasks, improves on video and spatial reasoning, and reaches up to 3.5x faster wall-clock inference. HF Daily Papers' note
score 6