Megadose AI progress, ranked and analyzed.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model

· HF Daily Papers ·
Mage-VL cuts visual tokens by more than 75% by reading video through codec signals instead of uniform frame sampling.

The paper says its Mage-ViT tokenizer selects dynamic, high-entropy regions using motion vectors and residual energy across I- and P-frames. The model is built for real-time multimodal understanding, with a lightweight event gate feeding a causal decoder for streaming perception. In evaluations, Mage-VL-4B matches Qwen3-VL-4B on static tasks, improves on video and spatial reasoning, and reaches up to 3.5x faster wall-clock inference. HF Daily Papers' note

score 6

Categories: Research