VideoLoop: Looped Working Memory Against Semantic Thrashing in Long-Form Video Agents
VideoLoop keeps a video agent’s working memory bounded by rewriting it after each reasoning step, instead of letting old context pile up.
The paper argues that append-only memory can add new evidence but cannot clear noise without a rewrite operation. VideoLoop pairs an outer video-reasoning loop with an inner loop that retrieves prior artifacts from a filesystem and rewrites the active context. In experiments, it improves four LVLM backbones by an average of 4.2 points on VideoMME long. With Gemini 3.1 Pro, it reports 88.3% on VideoMME long, 88.8% on VideoMMMU, and 80.9% on LongVideoBench long. HF Daily Papers' note
The paper argues that append-only memory can add new evidence but cannot clear noise without a rewrite operation. VideoLoop pairs an outer video-reasoning loop with an inner loop that retrieves prior artifacts from a filesystem and rewrites the active context. In experiments, it improves four LVLM backbones by an average of 4.2 points on VideoMME long. With Gemini 3.1 Pro, it reports 88.3% on VideoMME long, 88.8% on VideoMMMU, and 80.9% on LongVideoBench long. HF Daily Papers' note
score 4