Megadose AI progress, ranked and analyzed.

ViTAMINS: An Empirical Study of Training Self-Supervised Vision Transformers with Synthetic Hard Negatives

· ArXiv · AI/CL/LG ·
Synthetic hard negatives boosted self-supervised ViT representations across several vision benchmarks.

The paper says ViTAMINS adds synthetic hard negatives to existing contrastive pretraining with simple changes. It reports gains on ImageNet, transfer learning, retrieval, copy detection, and image/video segmentation tasks. The authors claim the learned representations carry explicit semantic information and can improve classifier performance by up to 11.3% over baselines. They also say a ViT-B setup beats V-JEPA with ViT-L while using fewer resources. ArXiv · AI/CL/LG's note

score 5

Categories: Research