Megadose AI progress, ranked and analyzed.

Last Translation Benchmark

· HF Daily Papers ·
The paper proposes a live benchmark built from examples that leading translation models already fail.

Its authors argue that standard machine-translation benchmarks are nearing saturation and that automatic metrics can be unreliable, gameable, and hard to act on. The Last Translation Benchmark uses human-authored, peer-reviewed texts, images, audio, and videos designed to expose concrete failure cases. Each example includes handcrafted verification rules so future evaluations can check specific translation errors. LTBv1 includes accepted contributions before September 1, 2026, with later releases planned as new data comes in. HF Daily Papers' note

score 5

Categories: Research