Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation
The paper proposes a way to make disaggregated AI evals more reliable when some domains have too few labeled examples.
It treats an evaluation set as a finite population and estimates mean performance for each domain. The authors introduce prediction-powered smoothing, plus a taxonomy-aware version that borrows strength across related reporting groups. They also derive a design-based cross-validation score for choosing between direct and smoothed estimators. In their benchmark and deployed-agent traffic studies, the smoothed estimators improved point and interval estimates while keeping coverage near nominal. HF Daily Papers' note
It treats an evaluation set as a finite population and estimates mean performance for each domain. The authors introduce prediction-powered smoothing, plus a taxonomy-aware version that borrows strength across related reporting groups. They also derive a design-based cross-validation score for choosing between direct and smoothed estimators. In their benchmark and deployed-agent traffic studies, the smoothed estimators improved point and interval estimates while keeping coverage near nominal. HF Daily Papers' note
score 4