Beyond F1: Evaluating Coverage and Failure Recovery in AI Model Security Scanners
Static scanners are increasingly used to identify executable or otherwise unsafe content in machine- learning artifacts, yet conventional evaluation metrics characterize only cases where a scanner yields a usable security judgment. We evaluate ModelScan, ModelAudit, and Fickling using a controlled, artifact-backed benchmark on a synthetic corpus of 170 Pickle and PyTorch focused artifacts across 145 specimen families, 135 of which have binary security ground truth and 10 of which are intentionally malformed without labels. We explicitly distinguish non-N/A coverage, analysis completion, defini
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 52%Frontier AI labs still won’t say how they’d contain a rogue model →
- PossiblePossibly related (embedding) · 54%High validation accuracy can conceal production risk: Using SHAP to expose and block proxy bias at runtime [P] →
- PossiblePossibly related (embedding) · 52%Agentic AI and cybersecurity, the story so far →
- LinkedLinked via arxiv author · 85%Qianlong Lan →
“Beyond F1: Evaluating Coverage and Failure Recovery in AI Model Security Scanners”
- LinkedLinked via arxiv author · 85%Vinothini Pandurangan →
“Beyond F1: Evaluating Coverage and Failure Recovery in AI Model Security Scanners”
- LinkedLinked via arxiv author · 85%Anuj Kaul →
“Beyond F1: Evaluating Coverage and Failure Recovery in AI Model Security Scanners”
- LinkedLinked via arxiv author · 85%Indranil Sanyal →
“Beyond F1: Evaluating Coverage and Failure Recovery in AI Model Security Scanners”
