How Much is a Human Right Worth? ECtHR-NPD: A Benchmark for Predicting Non-Pecuniary Damage Awards
The benchmark tests whether models can predict ECtHR non-pecuniary damages when the court gives no fixed calculation rule.
ECtHR-NPD covers 14,575 cases with nominal-euro awards and chronological splits. The authors compare feature-based models, retrieval, fine-tuned encoders, prompted decoder LMs, and knowledge-augmented agents. More complex LM and agent approaches do not reliably beat the strongest feature-based baseline. All model families struggle with zero-award cases and high-award calibration, especially on the harder test view. ArXiv · AI/CL/LG's note
ECtHR-NPD covers 14,575 cases with nominal-euro awards and chronological splits. The authors compare feature-based models, retrieval, fine-tuned encoders, prompted decoder LMs, and knowledge-augmented agents. More complex LM and agent approaches do not reliably beat the strongest feature-based baseline. All model families struggle with zero-award cases and high-award calibration, especially on the harder test view. ArXiv · AI/CL/LG's note
score 4