What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models
Regulatory compliance monitoring in deployed language models is increasingly implemented as a legal and audit control, checking model outputs against written rules spanning data protection, healthcare, financial regulation, and platform policy. Such monitoring is meaningful only if a detector's verdict depends on the stated rule rather than on surface features of the scenario. We show this condition fails across the current class of compliance detectors, a failure we call rule blindness. Deleting, permuting, or substituting the governing rule leaves detection accuracy unchanged for every guard
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 48%Understanding Annotator Safety Policy with Interpretability - Apple Machine Learning Research →
- PossiblePossibly related (embedding) · 47%Competition Agencies Compute. The Law Still Assumes They Read. - Wolters Kluwer →
- PossiblePossibly related (embedding) · 46%Revolutionizing software quality: new study explores large language models' pioneering role in defect detection - EurekAlert! →
- PossiblePossibly related (embedding) · 46%The biggest surprise while building an AI verification system wasn't the AI. →
- LinkedLinked via arxiv author · 85%Saisab Sadhu →
“What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models”
- LinkedLinked via arxiv author · 85%Aadit Sengupta →
“What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models”
- LinkedLinked via arxiv author · 85%Vinay Kumar Sankarapu →
“What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models”
- LinkedLinked via arxiv author · 85%Pratinav Seth →
“What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models”
