Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation
The paper proposes a way to estimate AI performance in thinly sampled domains without relying only on each domain’s sparse labels.
The authors frame evaluation sets as finite populations and target point and interval estimates for each domain mean. Their prediction-powered smoothing methods borrow strength across domains, including through a reporting taxonomy. They also introduce an approximately unbiased design-based cross-validation score to choose between direct and smoothed estimators. In benchmark and deployed-agent settings with fully observed outcomes, the proposed estimators improved point and interval estimation and kept near-nominal coverage. ArXiv · AI/CL/LG's note
The authors frame evaluation sets as finite populations and target point and interval estimates for each domain mean. Their prediction-powered smoothing methods borrow strength across domains, including through a reporting taxonomy. They also introduce an approximately unbiased design-based cross-validation score to choose between direct and smoothed estimators. In benchmark and deployed-agent settings with fully observed outcomes, the proposed estimators improved point and interval estimation and kept near-nominal coverage. ArXiv · AI/CL/LG's note
score 4