Megadose AI progress, ranked and analyzed.

Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination

· HF Daily Papers ·
The paper argues that “zero gap” aggregate scores can hide failed contamination cleanup.

Its proposed SA-PPG metric compares sampled per-question solve probabilities against a clean model, then groups results by the clean model’s own solve probability. The authors say older mitigation methods can look restored under G-AP because over- and under-suppression cancel out. They introduce RailCap, which suppresses likely memorized greedy-token paths during generation instead of relying on a prior contamination estimate. Across tested contaminated models and benchmarks, the paper says RailCap produces the lowest SA-PPG. HF Daily Papers' note

score 5

Categories: Research