When Names Cross Scripts: A Source-Grounded Benchmark for Historical Entity Reconciliation in the Mongol World
Source-grounded evidence turned name reconciliation from a guessing problem into a near-resolved one.
Chen and Zhang introduce MHER, a benchmark for matching historical person-name attestations from the Mongol world across languages, scripts, and transcription traditions. Its core set has 396 balanced name-only pairs, with a stricter 160-pair subset built from mention-by-source evidence. Across five generative systems, adding correct source evidence improved test accuracy by 12.96 to 94.44 percentage points over names alone. In identical-name cases involving different people, models got 0 of 25 decisions right with names only, but 24 of 25 with source-grounded evidence. ArXiv · AI/CL/LG's note
Chen and Zhang introduce MHER, a benchmark for matching historical person-name attestations from the Mongol world across languages, scripts, and transcription traditions. Its core set has 396 balanced name-only pairs, with a stricter 160-pair subset built from mention-by-source evidence. Across five generative systems, adding correct source evidence improved test accuracy by 12.96 to 94.44 percentage points over names alone. In identical-name cases involving different people, models got 0 of 25 decisions right with names only, but 24 of 25 with source-grounded evidence. ArXiv · AI/CL/LG's note
score 4