Image Classifiers are Efficient Self-Supervised Video Representation Learners
VideoMSN turns pretrained image ViTs into video learners by arranging sampled frames as “super images.”
The paper says the method builds paired masked views, one spatial and one temporal, then aligns them with a shared ViT encoder and masked Siamese loss. It avoids reconstruction and heavy 3D video architectures. Starting from DINO-v3 and DeiT-v3 encoders, it reports state-of-the-art results on Kinetics-400, UCF101, and HMDB51 with far fewer video pretraining epochs than prior methods. ArXiv · AI/CL/LG's note
The paper says the method builds paired masked views, one spatial and one temporal, then aligns them with a shared ViT encoder and masked Siamese loss. It avoids reconstruction and heavy 3D video architectures. Starting from DINO-v3 and DeiT-v3 encoders, it reports state-of-the-art results on Kinetics-400, UCF101, and HMDB51 with far fewer video pretraining epochs than prior methods. ArXiv · AI/CL/LG's note
score 4