Megadose AI progress, ranked and analyzed.

Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding

· HF Daily Papers ·
Parallel Tube Decoding cuts video grounding latency by 79x on VidSTG while improving accuracy.

The paper targets spatio-temporal video grounding: finding when a described event happens and tracking the relevant entity through that span. Its method decodes temporal boundaries first, then generates spatial boxes in parallel instead of autoregressively across the tube. The authors say this removes token- and trajectory-level dependencies, holding sequential decoding to two rounds regardless of tube length. A compact 4B model reports favorable results on VidSTG and HC-STVG, with zero-shot transfer to related video grounding and tracking tasks. HF Daily Papers' note

score 5

Categories: Research