Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning
The paper’s claim is that video models can learn visual foresight during training without generating intermediate images at inference.
Internalized Visual Thinking trains on unlabeled videos by predicting future-frame embeddings alongside the text answer. At inference, it answers directly, avoiding the frame synthesis and re-encoding used in explicit Visual CoT. The authors report gains over direct-answer fine-tuning across six evaluation settings, with comparable or better results than Visual CoT. They also report more than a 5x reduction in average end-to-end latency. HF Daily Papers' note
Internalized Visual Thinking trains on unlabeled videos by predicting future-frame embeddings alongside the text answer. At inference, it answers directly, avoiding the frame synthesis and re-encoding used in explicit Visual CoT. The authors report gains over direct-answer fine-tuning across six evaluation settings, with comparable or better results than Visual CoT. They also report more than a 5x reduction in average end-to-end latency. HF Daily Papers' note
score 5