The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping
Gemini 3.6 Flash could handle some slow, persistent state changes, but failed to reliably count transient blinks at all.
The paper tests video event counting across 2,190 controlled clips, with executable traces used to check not just final answers but the events models claimed to see. Performance collapses as event count and frequency rise: in the high-count, high-frequency setting, only 0.2% of final counts are correct. Raising the sampling rate improves one task’s final accuracy, but the reported event sequence matches ground truth only 3.7% of the time. The authors argue that aggregate accuracy hides where temporal reasoning breaks. ArXiv · AI/CL/LG's note
The paper tests video event counting across 2,190 controlled clips, with executable traces used to check not just final answers but the events models claimed to see. Performance collapses as event count and frequency rise: in the high-count, high-frequency setting, only 0.2% of final counts are correct. Raising the sampling rate improves one task’s final accuracy, but the reported event sequence matches ground truth only 3.7% of the time. The authors argue that aggregate accuracy hides where temporal reasoning breaks. ArXiv · AI/CL/LG's note
score 5