TempCloze: Can Video-LLMs Identify the Missing Middle?
Video-LLMs mostly fail at placing events at the right moment, even when they recognize what should happen.
TempCloze tests models by showing the beginning and end of a video and asking them to choose the true missing middle from four clips. The benchmark uses 1,521 filtered videos, with distractors designed to reduce language and appearance shortcuts. Across 10 proprietary and 21 open-source Video-LLMs, alignment was the main weakness: models handled plausible content and local progression better than temporal fit. The paper also studies how order, context direction, visible span, frame density, and test-time scaling affect model choices. HF Daily Papers' note
TempCloze tests models by showing the beginning and end of a video and asking them to choose the true missing middle from four clips. The benchmark uses 1,521 filtered videos, with distractors designed to reduce language and appearance shortcuts. Across 10 proprietary and 21 open-source Video-LLMs, alignment was the main weakness: models handled plausible content and local progression better than temporal fit. The paper also studies how order, context direction, visible span, frame density, and test-time scaling affect model choices. HF Daily Papers' note
score 5