Megadose Built for builders and researchers.

Cross-Lingual Alignment for Decoder-Only Models using MoE Routers

· HF Daily Papers ·
The paper tests router-output alignment as a substitute for hidden-state contrastive alignment in decoder-only MoE LLMs.

The authors argue multilingual tokenization makes direct representation alignment harder in decoder-only models. Their method aligns mixture-of-experts router outputs, which can be pooled across tokens for sequence-level comparison. In controlled continual pre-training on four open-source MoEs, the added routing loss also aligns hidden representations across languages. The paper reports improved multilingual performance on its evaluation suite. HF Daily Papers' note

score 4

Categories: Research