Ladders in Chaos: When, How, (and Perhaps Why) Does Test-Time Scaling Improve LLM Machine Translation
Sequential test-time sampling gave translation models a higher ceiling, especially when the sample budget was small.
The paper compares sequential attempts, where later translations build on earlier ones, with parallel sampling plus reranking. Its manual analysis finds sequential scaling improves fluency and naturalness, but can hurt accuracy at larger inference budgets. The authors partly attribute the gains to the model seeing more target-side context, while noting sensitivity to how that context is constructed. ArXiv · AI/CL/LG's note
The paper compares sequential attempts, where later translations build on earlier ones, with parallel sampling plus reranking. Its manual analysis finds sequential scaling improves fluency and naturalness, but can hurt accuracy at larger inference budgets. The authors partly attribute the gains to the model seeing more target-side context, while noting sensitivity to how that context is constructed. ArXiv · AI/CL/LG's note
score 4