SMAT: Simple and Efficient Merge-Aware Training
SMAT trains experts to survive the merge step, not just their own task loss.
The paper frames common model-merging methods as scale, mask, and perturb operations on expert updates. SMAT simulates those merged-parameter conditions during training while still using one forward and one backward pass per step. Across four language and vision-language backbones, it reports mean gains of 1.07 to 2.16 points over the strongest baseline for each backbone across five merging methods. Training-time overhead is reported at under 2% versus standard fine-tuning. HF Daily Papers' note
The paper frames common model-merging methods as scale, mask, and perturb operations on expert updates. SMAT simulates those merged-parameter conditions during training while still using one forward and one backward pass per step. Across four language and vision-language backbones, it reports mean gains of 1.07 to 2.16 points over the strongest baseline for each backbone across five merging methods. Training-time overhead is reported at under 2% versus standard fine-tuning. HF Daily Papers' note
score 4