TRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding
TRACE argues that streaming-video models need to be judged by when and how they act, not just whether their final answers are right.
The paper introduces a benchmark that annotates when evidence becomes valid, how much visual history is used, and what should trigger a response. In tests on 1,240 records from 517 videos, eight public models or systems showed that similar QA accuracy can hide different completion rates, answer validity, and generation workload. For proactive settings, the authors separate quality from response delay, false alarms, and missed target windows. HF Daily Papers' note
The paper introduces a benchmark that annotates when evidence becomes valid, how much visual history is used, and what should trigger a response. In tests on 1,240 records from 517 videos, eight public models or systems showed that similar QA accuracy can hide different completion rates, answer validity, and generation workload. For proactive settings, the authors separate quality from response delay, false alarms, and missed target windows. HF Daily Papers' note
score 5