Cross-Lingual Alignment for Decoder-Only Models using MoE Routers
The paper tests router-output alignment as a substitute for hidden-state contrastive alignment in decoder-only MoE LLMs.
The authors argue multilingual tokenization makes direct representation alignment harder in decoder-only models. Their method aligns mixture-of-experts router outputs, which can be pooled across tokens for sequence-level comparison. In controlled continual pre-training on four open-source MoEs, the added routing loss also aligns hidden representations across languages. The paper reports improved multilingual performance on its evaluation suite. HF Daily Papers' note
The authors argue multilingual tokenization makes direct representation alignment harder in decoder-only models. Their method aligns mixture-of-experts router outputs, which can be pooled across tokens for sequence-level comparison. In controlled continual pre-training on four open-source MoEs, the added routing loss also aligns hidden representations across languages. The paper reports improved multilingual performance on its evaluation suite. HF Daily Papers' note
score 4