TempCloze: Can Video-LLMs Identify the Missing Middle?
The benchmark tests whether video models can pick the true missing middle between a clip’s beginning and end.
TempCloze uses 1,521 filtered videos, mostly long-take and egocentric, with four candidate middle clips per prompt. Its distractors are built to separate semantic fit, timing alignment, and event progression while reducing easy appearance cues. In tests across 10 proprietary and 21 open-source Video-LLMs, the main failure point was temporal alignment: models could often spot plausible events, but not when they belonged. The authors also probe how order, context direction, visible span, frame density, and test-time scaling affect model choices. Source: ArXiv · AI/CL/LG's note
TempCloze uses 1,521 filtered videos, mostly long-take and egocentric, with four candidate middle clips per prompt. Its distractors are built to separate semantic fit, timing alignment, and event progression while reducing easy appearance cues. In tests across 10 proprietary and 21 open-source Video-LLMs, the main failure point was temporal alignment: models could often spot plausible events, but not when they belonged. The authors also probe how order, context direction, visible span, frame density, and test-time scaling affect model choices. Source: ArXiv · AI/CL/LG's note
score 5