Megadose Built for builders and researchers.

Before It Fades: Reinforcing Temporal Representations at Inference Time in VideoLLMs

· HF Daily Papers ·
VideoLLMs appear to capture temporal order mid-network, then lose that signal before the final answer.

The paper tests this by reversing video frames and measuring how representations change layer by layer. It finds temporal divergence peaks in intermediate layers and fades toward the output, which can leave reversed videos with unchanged predictions. The proposed fix, Temporal Activation Injection, reuses that peak temporal signal in later layers at inference time, without training. The authors report gains across three VideoLLMs and four benchmarks, with little effect on non-temporal tasks. HF Daily Papers' note

score 4

Categories: Research