What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models
Tested compliance guards often ignored the governing rule and classified from scenario cues. Builders should counterfactually swap policies before trusting a detector as an audit control.
Across every tested guard and activation probe, deleting, permuting, or replacing the governing rule left accuracy unchanged. Even a policy-conditioned guard could cite a clause while barely changing its verdict when the clause became permissive; the authors call this **rule blindness**.
Audit compliance detectors with crossed examples where neither the scenario nor rule alone predicts the label. The paper found **step-by-step reasoning** escaped the failure seen in fast detectors and released a counterfactual protocol for testing future claims.
Across every tested guard and activation probe, deleting, permuting, or replacing the governing rule left accuracy unchanged. Even a policy-conditioned guard could cite a clause while barely changing its verdict when the clause became permissive; the authors call this **rule blindness**. Audit compliance detectors with crossed examples where neither the scenario nor rule alone predicts the label. The paper found **step-by-step reasoning** escaped the failure seen in fast detectors and released a counterfactual protocol for testing future claims. The proposed **Internal Compliance Score** missed its preregistered bar, matched a bag-of-words baseline in pooled generalization, and lost its ranking gain under an adaptive white-box attack. Its advantage is low-cost auditing, not demonstrated rule understanding.
This turns general warnings about weak verifiers into a specific audit requirement: vary the governing rule independently of the scenario and verify that verdicts follow the rule. It also sharply limits activation-probe and fast-guard claims, because tested accuracy may reflect scenario cues rather than compliance reasoning; step-by-step reasoning is the reported exception.