TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs
TimeLens2 is built to mark the exact video intervals that support an answer, including multiple spans.
The paper frames temporal grounding as a set-valued task rather than a single timestamp or segment prediction problem. Its TimeLens2-93K training data uses proposal generation, independent localization, consensus, semantic checks, and boundary refinement to improve multi-span labels. The authors also introduce a temporal Wasserstein reward, paired with temporal IoU, to give feedback without brittle segment matching. Across seven benchmarks, the 2B model beats size-matched baselines, while 4B and 8B variants are reported as state of the art. HF Daily Papers' note
The paper frames temporal grounding as a set-valued task rather than a single timestamp or segment prediction problem. Its TimeLens2-93K training data uses proposal generation, independent localization, consensus, semantic checks, and boundary refinement to improve multi-span labels. The authors also introduce a temporal Wasserstein reward, paired with temporal IoU, to give feedback without brittle segment matching. Across seven benchmarks, the 2B model beats size-matched baselines, while 4B and 8B variants are reported as state of the art. HF Daily Papers' note
score 5