Megadose Built for builders and researchers.

The Exceedance Design Effect: Effective Sample Size for Thresholds under Clustering

· ArXiv · AI/CL/LG ·
Clustered calibration data can make 25,028 examples behave like about 1,300 for threshold reliability.

The paper argues that quantile thresholds in conformal prediction, abstention, and safety filtering need their own effective sample size when examples are clustered. That count depends on whether clustered scores fall on the same side of the chosen threshold, not on numeric score similarity. It says the correction currently used in conformal literature can be wrong in either direction, and that a dataset has a different effective sample size at each threshold level. The risk may not show up in average coverage across repeated runs, but it affects a single deployed system. ArXiv · AI/CL/LG's note

score 4

Categories: Research