One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts
A recurrent ViT using one shared block claims full-depth encoder accuracy with far fewer stored parameters.
The paper’s reViT repeats a single Transformer block and varies the FFN by mixing a small bank of shared experts according to depth. The authors report that this restores depth-specific behavior without intermediate feature distillation and works under both ImageNet-1k training and DINOv2 distillation. In their tests, weight-space merging beat token-dispatch and output-mixture MoE alternatives at the same one-FFN budget. A reViT-B/16 trained from scratch matched DeiT III accuracy with about 70% fewer stored parameters, while an 8-expert distilled model kept nearly all of its DINOv2 teacher’s linear-probe accuracy. ArXiv · AI/CL/LG's note
The paper’s reViT repeats a single Transformer block and varies the FFN by mixing a small bank of shared experts according to depth. The authors report that this restores depth-specific behavior without intermediate feature distillation and works under both ImageNet-1k training and DINOv2 distillation. In their tests, weight-space merging beat token-dispatch and output-mixture MoE alternatives at the same one-FFN budget. A reViT-B/16 trained from scratch matched DeiT III accuracy with about 70% fewer stored parameters, while an 8-expert distilled model kept nearly all of its DINOv2 teacher’s linear-probe accuracy. ArXiv · AI/CL/LG's note
score 4