PaMIR: Open Benchmark of Public Credit-Default Datasets
PaMIR packages 19 public credit-default datasets into a leakage-audited benchmark for scarce, delayed labels.
The release covers 1.24 million loans, firms, and card accounts from nine countries, rebuilt from pinned source snapshots rather than redistributed. Models are evaluated as single functions under both repeated i.i.d. splits and a label-delayed stream where applications are scored on arrival. AUC is reported by label budget, and aggregate fleet means are withheld unless every dataset is scored. The paper also includes a synthetic-data harness that tests generated training rows without exposing held-out rows to the generator. ArXiv · AI/CL/LG's note
The release covers 1.24 million loans, firms, and card accounts from nine countries, rebuilt from pinned source snapshots rather than redistributed. Models are evaluated as single functions under both repeated i.i.d. splits and a label-delayed stream where applications are scored on arrival. AUC is reported by label budget, and aggregate fleet means are withheld unless every dataset is scored. The paper also includes a synthetic-data harness that tests generated training rows without exposing held-out rows to the generator. ArXiv · AI/CL/LG's note
score 4