Megadose Built for builders and researchers.

What's the Catch? Evaluating Temporal Consistency in Vision-Language Models

· ArXiv · AI/CL/LG ·
VLMs caught noisy frames, but mostly missed swapped-frame time errors.

The paper introduces TimeCatch, a benchmark that tests temporal grounding by inserting anomalies into frame sequences. Models detected and often localized frame-level Gaussian-noise anomalies, but performed near chance when consecutive frames were swapped. Humans were near ceiling on both detection and localization. The authors say the gap persists across model scale, prompting, sequence length, and visual similarity, pointing to weak cross-frame temporal reasoning. ArXiv · AI/CL/LG's note

score 4

Categories: Research