NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap
The benchmark finds the English-Korean gap is not mainly about Korean wording, but about tasks that require sub-syllabic Hangul execution.
NOLLI tests 15 puzzle types across 7,500 seed-regenerable items, with difficulty calibrated by model behavior rather than puzzle size. Across models above the paper’s accuracy floor, matched English and Korean translations were statistically equivalent within a 10-point margin. The sharpest split appears in writing-system-heavy tasks: Korean Cipher trails English by as much as 68.7 points, while jamo-based cryptarithmetic does not show the same penalty. The authors frame the result as diagnostic, consistent with difficulty in multi-step work over Hangul jamo, not proof of a single cause. HF Daily Papers' note
NOLLI tests 15 puzzle types across 7,500 seed-regenerable items, with difficulty calibrated by model behavior rather than puzzle size. Across models above the paper’s accuracy floor, matched English and Korean translations were statistically equivalent within a 10-point margin. The sharpest split appears in writing-system-heavy tasks: Korean Cipher trails English by as much as 68.7 points, while jamo-based cryptarithmetic does not show the same penalty. The authors frame the result as diagnostic, consistent with difficulty in multi-step work over Hangul jamo, not proof of a single cause. HF Daily Papers' note
score 4