Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies
GEB turns recurring objects in long videos into grounded “biographies” the model can retrieve.
The paper argues that timelines and text labels miss when the same physical object reappears across hours or days. Its framework groups visually grounded observations of one instance across clips, then retrieves that biography alongside event evidence for question answering. Across four benchmarks, including day- and week-long recordings, it beats prior memory systems; on EgoLifeQA it reports 72.0% accuracy, 4.4 points above the best published result. Ablations say the gains come from identity association and biography reading, not just adding more descriptions. ArXiv · AI/CL/LG's note
The paper argues that timelines and text labels miss when the same physical object reappears across hours or days. Its framework groups visually grounded observations of one instance across clips, then retrieves that biography alongside event evidence for question answering. Across four benchmarks, including day- and week-long recordings, it beats prior memory systems; on EgoLifeQA it reports 72.0% accuracy, 4.4 points above the best published result. Ablations say the gains come from identity association and biography reading, not just adding more descriptions. ArXiv · AI/CL/LG's note
score 4