Megadose AI progress, ranked and analyzed.

What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models

· ArXiv · AI/CL/LG ·
The paper says tested compliance guards often ignore the rule they are supposed to enforce.

The authors call the failure “rule blindness”: deleting, shuffling, or swapping the governing rule did not change accuracy for the guards and activation probes they tested. One policy-conditioned guard could cite the relevant clause while barely changing its verdict when that clause was replaced with a permissive one. Their proposed Internal Compliance Score was cheap enough for audits, but it did not beat their registered baseline test and matched a bag-of-words model on pooled generalization. ArXiv · AI/CL/LG's note

score 4

Categories: Research