Megadose AI progress, ranked and analyzed.

I Have a Stream: Making Self-Supervised Learning Work on Continuous Video

· HF Daily Papers ·
The paper finds that near-duplicate frames inside each streaming batch are the main obstacle to self-supervised video pretraining.

The authors test strict temporal training on WT++, a 95-hour urban walking-tour video dataset, without reshuffling or replaying frames across epochs. Contrastive and distillation methods struggle under that setup, while MAE holds up better but still trails standard i.i.d. pretraining. Their StreamMAE keeps the MAE reconstruction objective and changes the input pipeline with stream-aware regularization and motion-biased crop selection. It beats streaming baselines, matches i.i.d. MAE on the same video data, and improves as the stream grows from 12 to 95 hours. HF Daily Papers' note

score 4

Categories: Research