Time-Reversed Imaging: A Multimodal Benchmark and Framework for Reconstructing Past Human-Environment Interactions
The paper tests whether recent actions can be reconstructed from residual thermal, ultraviolet, and visible traces.
The authors introduce TRACE-HEI, a proof-of-concept dataset of tri-modal video after actions including sitting, touching, moving objects, and liquid spills. Recordings cover different materials and extend up to three minutes after contact. Their benchmark turns detected traces into structured text, then uses that to guide a vision-language diffusion model toward plausible past frames. The abstract says the task remains difficult, but complementary modalities reduce ambiguity enough to make it feasible. ArXiv · AI/CL/LG's note
The authors introduce TRACE-HEI, a proof-of-concept dataset of tri-modal video after actions including sitting, touching, moving objects, and liquid spills. Recordings cover different materials and extend up to three minutes after contact. Their benchmark turns detected traces into structured text, then uses that to guide a vision-language diffusion model toward plausible past frames. The abstract says the task remains difficult, but complementary modalities reduce ambiguity enough to make it feasible. ArXiv · AI/CL/LG's note
score 5