Last Translation Benchmark
The paper proposes a live benchmark built from examples that leading translation models still fail.
The authors argue that standard machine-translation benchmarks are nearing saturation and that automatic metrics can be unreliable or gamed. Their dataset includes human-authored, peer-reviewed text, image, audio, and video examples designed to expose concrete failure cases. Each example comes with handcrafted verification rules so future evaluations can identify what broke. LTBv1 covers accepted contributions before September 1, 2026, with more releases planned as new data is added. ArXiv · AI/CL/LG's note
The authors argue that standard machine-translation benchmarks are nearing saturation and that automatic metrics can be unreliable or gamed. Their dataset includes human-authored, peer-reviewed text, image, audio, and video examples designed to expose concrete failure cases. Each example comes with handcrafted verification rules so future evaluations can identify what broke. LTBv1 covers accepted contributions before September 1, 2026, with more releases planned as new data is added. ArXiv · AI/CL/LG's note
score 5