Megadose Built for builders and researchers.

One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts

· HF Daily Papers ·
reViT claims a recurrent single-block vision encoder can approach full-depth performance while storing far fewer parameters.

The paper uses a shared Transformer block repeatedly, with depth-specific FFN behavior produced by mixtures from a small expert bank.
In its tests, weight-space merging was the strongest MoE variant under the same one-FFN budget.
Trained from scratch, reViT-B/16 matched DeiT III accuracy with about 70% fewer stored parameters.
With DINOv2 distillation, an 8-expert model retained nearly all teacher linear-probe accuracy and transferred to classification, segmentation, and depth prediction.
Source: HF Daily Papers' note

score 5

Categories: Research