Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination
The paper argues that a zero aggregate gap can still hide failed contamination cleanup.
The authors say G-AP can miss per-question damage because errors cancel when scores are averaged. They propose SA-PPG, which compares sampled solve probabilities question by question against a clean model and groups results by the clean model’s probability bands. Their mitigation method, RailCap, suppresses tokens when generation falls back onto a greedy memorized path. Across the tested contaminated models and benchmarks, they report that older mitigation methods look better under G-AP than they do under SA-PPG, while RailCap gets the lowest SA-PPG. ArXiv · AI/CL/LG's note
The authors say G-AP can miss per-question damage because errors cancel when scores are averaged. They propose SA-PPG, which compares sampled solve probabilities question by question against a clean model and groups results by the clean model’s probability bands. Their mitigation method, RailCap, suppresses tokens when generation falls back onto a greedy memorized path. Across the tested contaminated models and benchmarks, they report that older mitigation methods look better under G-AP than they do under SA-PPG, while RailCap gets the lowest SA-PPG. ArXiv · AI/CL/LG's note
score 4