Megadose Built for builders and researchers.

Foresight: planning future perception in streaming VLMs without retraining

· HF Daily Papers ·
FORESIGHT lets a frozen streaming VLM plan where to spend future compute while the video is still arriving.

The paper says streaming VLMs can anticipate near-future context well enough to guide inference without retraining. FORESIGHT runs a second, weight-shared LLM ahead of the live stream to decide when to reason next, what to check, and how densely to sample. Using a frozen Qwen3-VL-8B backbone, it reports 23.0 mean joint F1 on OmniPro Online, 9.5% above the strongest trained baseline. The authors say the biggest gain comes when key evidence appears later in the video stream. HF Daily Papers' note

score 4

Categories: Research