What AI Red-Team Evaluations Can and Cannot Prove
Red-team results can only certify safety inside a computable evidentiary boundary.
Bandana Kaur defines an “evidential ceiling” for AI red-team evaluations: the maximum amount a test result can shift belief under a fixed testing budget. The paper argues modest benchmarks can support claims about high-frequency harms, but feasible passive benchmarks fall far short for rare catastrophic harms. The same bound is framed to cover adaptive and automated red teaming, where the key issue is discrimination between hypotheses rather than attack success alone. Version 2 corrects figures, sample-size calculations, and captions without changing the paper’s conclusions. HF Daily Papers' note
Bandana Kaur defines an “evidential ceiling” for AI red-team evaluations: the maximum amount a test result can shift belief under a fixed testing budget. The paper argues modest benchmarks can support claims about high-frequency harms, but feasible passive benchmarks fall far short for rare catastrophic harms. The same bound is framed to cover adaptive and automated red teaming, where the key issue is discrimination between hypotheses rather than attack success alone. Version 2 corrects figures, sample-size calculations, and captions without changing the paper’s conclusions. HF Daily Papers' note
score 5