Sign InOpen Brain
arXivPaperNeeds Review

What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models

Tested compliance guards often ignored the governing rule and classified from scenario cues. Builders should counterfactually swap policies before trusting a detector as an audit control.

arXiv · Aug 17, 2026
Open Source Open MarkdownOpen JSON
Source Summary

Across every tested guard and activation probe, deleting, permuting, or replacing the governing rule left accuracy unchanged. Even a policy-conditioned guard could cite a clause while barely changing its verdict when the clause became permissive; the authors call this **rule blindness**.

Practical Implication

Audit compliance detectors with crossed examples where neither the scenario nor rule alone predicts the label. The paper found **step-by-step reasoning** escaped the failure seen in fast detectors and released a counterfactual protocol for testing future claims.

Agent-Ready Context
Across every tested guard and activation probe, deleting, permuting, or replacing the governing rule left accuracy unchanged. Even a policy-conditioned guard could cite a clause while barely changing its verdict when the clause became permissive; the authors call this **rule blindness**.

Audit compliance detectors with crossed examples where neither the scenario nor rule alone predicts the label. The paper found **step-by-step reasoning** escaped the failure seen in fast detectors and released a counterfactual protocol for testing future claims.

The proposed **Internal Compliance Score** missed its preregistered bar, matched a bag-of-words baseline in pooled generalization, and lost its ranking gain under an adaptive white-box attack. Its advantage is low-cost auditing, not demonstrated rule understanding.
Connected Context · Feed7 Judgment

This turns general warnings about weak verifiers into a specific audit requirement: vary the governing rule independently of the scenario and verify that verdicts follow the rule. It also sharply limits activation-probe and fast-guard claims, because tested accuracy may reflect scenario cues rather than compliance reasoning; step-by-step reasoning is the reported exception.

Form, Not Content? A Preregistered, Placebo-Controlled Evaluation of Learned Error-Conditioned Self-Repair Through Prompts and Weights in Frozen Small Code ModelsBoth use controlled substitutions to test whether a system responds to the claimed content; each finds that apparent success can persist when the supposedly causal information is replaced or mismatched.When Will The Benchmaxxing Plague End? — Nick Heiner, Surge AIThe paper supplies a concrete instance of the weak-verifier and shortcut problem: compliance scores remain stable when the governing policy changes.Cursor earns AIUC-1 certification for agent security and reliabilityAIUC-1 adds live adversarial testing to controls audits; this paper identifies crossed policy–scenario counterfactuals as a useful test for whether compliance guards actually enforce those controls.MakazhanAlpamys/SoupSoup’s over-refusal and noise-floor checks already argue against trusting one aggregate score; rule-blind detectors add a distinct need to test whether the intended rule causally changes verdicts.
Context Map
benchmarksecurity#agent-evals#benchmark-integrity#agent-reliability
Uncertainty
The proposed **Internal Compliance Score** missed its preregistered bar, matched a bag-of-words baseline in pooled generalization, and lost its ranking gain under an adaptive white-box attack. Its advantage is low-cost auditing, not demonstrated rule understanding.