Too Sure to Be Safe: Model Calibration for Reliable Log Anomaly Detection
Online log anomaly detection is critical for maintaining the reliability of large-scale computing systems. Although recent language model-based log anomaly detectors achieve strong detection performance, their confidence estimates remain poorly calibrated. We show that these detectors frequently assign excessive confidence to incorrect predictions, particularly for anomalous logs under severe class imbalance. Moreover, confidence on erroneous predictions remains persistently high even when conventional calibration metrics indicate good calibration, creating a critical reliability gap for opera
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 48%Something keeps turning up in my prompt injection detection logs that I didn't expect. Curious if others doing LLM security work have seen it. [D] →
- PossiblePossibly related (embedding) · 46%Revolutionizing software quality: new study explores large language models' pioneering role in defect detection - EurekAlert! →
- FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- LinkedLinked via arxiv author · 85%Yongbin Li →
“Too Sure to Be Safe: Model Calibration for Reliable Log Anomaly Detection”
- LinkedLinked via arxiv author · 85%Dongdong Wang →
“Too Sure to Be Safe: Model Calibration for Reliable Log Anomaly Detection”
- LinkedLinked via arxiv author · 85%Siyang Lu →
“Too Sure to Be Safe: Model Calibration for Reliable Log Anomaly Detection”
