The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems
Current AI safety discourse still focuses disproportionately on visible failures, including obvious harms, dramatic misuse, and hypothetical catastrophic scenarios. That focus is incomplete. In deployed systems, many of the most consequential failures are quieter: plausible rather than spectacular, distributed across components rather than localized in a single output, and normalized by workflows before they are recognized as hazards. We argue that a central safety challenge in modern AI systems is increasingly not only whether a model emits a harmful response, but whether the broader socio-te
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 70%Safety and alignment in an era of long-horizon models →
- PossiblePossibly related (embedding) · 65%Nemotron 3.5 Content Safety: Customizable Multimodal Safety for Global Enterprise AI →
- PossiblePossibly related (embedding) · 64%Verisight →
- PossiblePossibly related (embedding) · 63%New Research: AI models can give dangerous responses despite output guardrails - The AI Journal →
- PossiblePossibly related (embedding) · 62%"Dangerous" AI models are coming no matter what →
- LinkedLinked via arxiv author · 85%Gjergji Kasneci →
“The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems”
- LinkedLinked via arxiv author · 85%Enkelejda Kasneci →
“The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems”
