Foresight: planning future perception in streaming VLMs without retraining
FORESIGHT lets a frozen streaming VLM plan where to spend future compute while the video is still arriving.
The paper says streaming VLMs can anticipate near-future context well enough to guide inference without retraining. FORESIGHT runs a second, weight-shared LLM ahead of the live stream to decide when to reason next, what to check, and how densely to sample. Using a frozen Qwen3-VL-8B backbone, it reports 23.0 mean joint F1 on OmniPro Online, 9.5% above the strongest trained baseline. The authors say the biggest gain comes when key evidence appears later in the video stream. HF Daily Papers' note
The paper says streaming VLMs can anticipate near-future context well enough to guide inference without retraining. FORESIGHT runs a second, weight-shared LLM ahead of the live stream to decide when to reason next, what to check, and how densely to sample. Using a frozen Qwen3-VL-8B backbone, it reports 23.0 mean joint F1 on OmniPro Online, 9.5% above the strongest trained baseline. The authors say the biggest gain comes when key evidence appears later in the video stream. HF Daily Papers' note
score 4