Megadose Built for builders and researchers.

LVMT: Video Mask Transformer for Long-term Video Segmentation

· HF Daily Papers ·
LVMT targets long video segmentation by adding learned memory selection and chunked long-video training.

The paper says current online methods lose objects in complex videos with long occlusions. LVMT uses a lightweight GRU-based propagation module to decide what object information stays in memory over time. Its Truncated Query Propagation trains on frame chunks while carrying object information between chunks, avoiding out-of-memory issues and vanishing gradients. Across six benchmarks, the authors report state-of-the-art results and 10X speed over the prior state of the art. HF Daily Papers' note

score 4

Categories: Research