Copy the Same, Distill the Difference: Initializing Linear Vision Transformers
The paper says Softmax ViT weights can help linear ViTs, but only if MLPs are copied and attention behavior is distilled.
The authors find that directly copying attention weights across the Softmax-to-linear boundary barely helps and can underperform random initialization. MLP weights transfer much better because they carry learned representations in a way the paper treats as operator-agnostic. With copied MLPs plus attention distillation, linear ViTs can close the gap with Softmax models and sometimes surpass them across tested variants, sizes, and datasets. ArXiv · AI/CL/LG's note
The authors find that directly copying attention weights across the Softmax-to-linear boundary barely helps and can underperform random initialization. MLP weights transfer much better because they carry learned representations in a way the paper treats as operator-agnostic. With copied MLPs plus attention distillation, linear ViTs can close the gap with Softmax models and sometimes surpass them across tested variants, sizes, and datasets. ArXiv · AI/CL/LG's note
score 4