CADER: Confidence-Aware Dynamic Evidence Reasoning for Long-Video Understanding
CADER decides at inference time whether a video question needs extra evidence work.
The framework starts with a global pass over sampled frames and uses a logit-margin confidence signal to let easier cases exit early. When confidence is low, it runs a second stage with temporal cropping, semantic verification, and relevance-guided resampling to find question-relevant moments. The authors report gains across multiple VideoQA benchmarks while skipping the heavier stage for high-confidence samples. ArXiv · AI/CL/LG's note
The framework starts with a global pass over sampled frames and uses a logit-margin confidence signal to let easier cases exit early. When confidence is low, it runs a second stage with temporal cropping, semantic verification, and relevance-guided resampling to find question-relevant moments. The authors report gains across multiple VideoQA benchmarks while skipping the heavier stage for high-confidence samples. ArXiv · AI/CL/LG's note
score 5