Presentation, Not Mechanism: A Render Confound in Deprecation-Aware Memory Evaluation
The reported RevisionLedger advantage mostly vanishes when the paper controls for how evidence is rendered.
The authors test deprecation-aware memory on 2,907 questions drawn from GitHub histories, Wikipedia, and temporal streams. In restored-value cases, RevisionLedger’s apparent +0.182 gain over flat retrieval is attributed almost entirely to easier presentation. With the render held fixed, its mechanism-only residual is near zero, while coarse invalidation performs best for current-state queries. The paper’s practical recommendation is to keep evaluation layouts matched and use the coarsest retained state sufficient for the query. ArXiv · AI/CL/LG's note
The authors test deprecation-aware memory on 2,907 questions drawn from GitHub histories, Wikipedia, and temporal streams. In restored-value cases, RevisionLedger’s apparent +0.182 gain over flat retrieval is attributed almost entirely to easier presentation. With the render held fixed, its mechanism-only residual is near zero, while coarse invalidation performs best for current-state queries. The paper’s practical recommendation is to keep evaluation layouts matched and use the coarsest retained state sufficient for the query. ArXiv · AI/CL/LG's note
score 5