Megadose AI progress, ranked and analyzed.

Partially Correlated Verifier Cascades in LLM Harnesses: Concave Log-Odds, Polynomial Reliability, and Blind-Spot Ceilings

· ArXiv · AI/CL/LG ·
Verifier cascades can hit a hard reliability ceiling when their mistakes are correlated.

Jiangang Han models false accepts as a latent per-instance rate, so repeated gates increasingly filter for errors the verifiers are already prone to miss. In that setting, log-odds gains are concave, not linear, and Beta latent models make failure decay polynomially rather than exponentially. A blind-spot mass at perfect false-accept rate caps how much evidence any number of gates can extract. Synthetic tests in the note show independence-based extrapolation understating failure by 20x at five gates and about 3000x at ten. ArXiv · AI/CL/LG's note

score 4

Categories: Research