Sound Probabilistic Safety Bounds for Large Language Models
The paper claims a way to put formally sound lower bounds on an LLM’s probability of producing harmful output.
The authors apply Clopper-Pearson confidence intervals to compute PAC-style safety bounds for a given prompt. Their algorithm searches the autoregressive generation tree by using latent-space features to prioritize branches more likely to turn harmful. They say this can produce useful lower bounds even when the true harm probability is very small. ArXiv · AI/CL/LG's note
The authors apply Clopper-Pearson confidence intervals to compute PAC-style safety bounds for a given prompt. Their algorithm searches the autoregressive generation tree by using latent-space features to prioritize branches more likely to turn harmful. They say this can produce useful lower bounds even when the true harm probability is very small. ArXiv · AI/CL/LG's note
score 5