Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation
MiLMMT-46-v1.0 improves open-model translation across 46 languages using reward-based post-training without reference translations.
The paper starts from supervised-finetuned MiLMMT-46 models, then applies GRPO with reference-free quality-estimation rewards gated by language ID. The authors combine the SFT and RL checkpoints through linear interpolation for the released v1.0 models. They report gains over the SFT versions, stronger recent open baselines, and leading reference-free scores against evaluated proprietary systems. They also test on-policy distillation, which matches but does not exceed the RL-plus-interpolation frontier. HF Daily Papers' note
The paper starts from supervised-finetuned MiLMMT-46 models, then applies GRPO with reference-free quality-estimation rewards gated by language ID. The authors combine the SFT and RL checkpoints through linear interpolation for the released v1.0 models. They report gains over the SFT versions, stronger recent open baselines, and leading reference-free scores against evaluated proprietary systems. They also test on-policy distillation, which matches but does not exceed the RL-plus-interpolation frontier. HF Daily Papers' note
score 6