Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination
The paper argues that “zero gap” aggregate scores can hide failed contamination cleanup.
Its proposed SA-PPG metric compares sampled per-question solve probabilities against a clean model, then groups results by the clean model’s own solve probability. The authors say older mitigation methods can look restored under G-AP because over- and under-suppression cancel out. They introduce RailCap, which suppresses likely memorized greedy-token paths during generation instead of relying on a prior contamination estimate. Across tested contaminated models and benchmarks, the paper says RailCap produces the lowest SA-PPG. HF Daily Papers' note
Its proposed SA-PPG metric compares sampled per-question solve probabilities against a clean model, then groups results by the clean model’s own solve probability. The authors say older mitigation methods can look restored under G-AP because over- and under-suppression cancel out. They introduce RailCap, which suppresses likely memorized greedy-token paths during generation instead of relying on a prior contamination estimate. Across tested contaminated models and benchmarks, the paper says RailCap produces the lowest SA-PPG. HF Daily Papers' note
score 5