Megadose AI progress, ranked and analyzed.

Layered LLM Defenses as an Ensemble: Access Tiers, Inference Cost, and the Measured Failure Correlation Between Defense Layers

· ArXiv · AI/CL/LG ·
The paper finds stacked LLM defenses failed together often enough to blunt the expected ensemble gains.

In a seven-layer defense stack, every measurable pair showed positive failure correlation, with phi values from 0.30 to 0.75. The measured joint residual attack success exceeded the independence-based prediction by as much as 0.172. The stack also refused four in five benign prompts while performing statistically no better than its strongest single layer. The authors argue the dependence comes from shared architecture around the same model, so stack performance has to be measured end to end. ArXiv · AI/CL/LG's note

score 5

Categories: Research