LVMT: Video Mask Transformer for Long-term Video Segmentation
LVMT targets long video segmentation by adding learned memory selection and chunked long-video training.
The paper says current online methods lose objects in complex videos with long occlusions. LVMT uses a lightweight GRU-based propagation module to decide what object information stays in memory over time. Its Truncated Query Propagation trains on frame chunks while carrying object information between chunks, avoiding out-of-memory issues and vanishing gradients. Across six benchmarks, the authors report state-of-the-art results and 10X speed over the prior state of the art. HF Daily Papers' note
The paper says current online methods lose objects in complex videos with long occlusions. LVMT uses a lightweight GRU-based propagation module to decide what object information stays in memory over time. Its Truncated Query Propagation trains on frame chunks while carrying object information between chunks, avoiding out-of-memory issues and vanishing gradients. Across six benchmarks, the authors report state-of-the-art results and 10X speed over the prior state of the art. HF Daily Papers' note
score 4