Length-Adaptive Decoding for Masked Diffusion Machine Translation
The paper argues that masked diffusion translation gains more from choosing the right output length than from changing token reveal order.
Its Entropy-Valley method picks a target canvas by scoring candidate lengths with predictive entropy from all-mask passes, without training changes. Against a corpus-statistics length baseline, it recovers large portions of the COMET-22 gain seen when reference target lengths are supplied. Expert evaluation supports adequacy gains for En-Zh, with stronger evidence on Zh-En. The paper was accepted to EMNLP 2026. HF Daily Papers' note
Its Entropy-Valley method picks a target canvas by scoring candidate lengths with predictive entropy from all-mask passes, without training changes. Against a corpus-statistics length baseline, it recovers large portions of the COMET-22 gain seen when reference target lengths are supplied. Expert evaluation supports adequacy gains for En-Zh, with stronger evidence on Zh-En. The paper was accepted to EMNLP 2026. HF Daily Papers' note
score 4