Context-structured Video Anomaly Detection with Large Vision-Language Models
Training video anomaly detectors is challenging due to the difficulty and cost of annotating diverse and rare abnormal events. Although recent large vision-language models enable training-free inference, existing approaches mostly rely on holistic inference over sampled video and may miss context-specific anomaly cues. In this paper, we present CSI-VAD, a training-free video anomaly detector that identifies abnormal events across diverse contexts. The key idea is to decompose each video into three distinct contexts (environment, objects, time) and perform context-specific inference in separate
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 59%VioletVision-3B →
“Fuzzy title match (0.73): “Context-structured Video Anomaly Detection with Large Vision” ≈ “VioletVision-3B””
- FuzzySimilar title/name (fuzzy) · 84%pytorch/vision →
“Fuzzy title match (0.92): “Context-structured Video Anomaly Detection with Large Vision” ≈ “pytorch/vision””
- FuzzySimilar title/name (fuzzy) · 59%Developer-Y/cs-video-courses →
“Fuzzy title match (0.73): “Context-structured Video Anomaly Detection with Large Vision” ≈ “Developer-Y/cs-video-courses””
- LinkedLinked via arxiv author · 85%Dongjun Kim →
“Context-structured Video Anomaly Detection with Large Vision-Language Models”
- LinkedLinked via arxiv author · 85%Changjae Oh →
“Context-structured Video Anomaly Detection with Large Vision-Language Models”
- LinkedLinked via arxiv author · 85%Andrea Cavallaro →
“Context-structured Video Anomaly Detection with Large Vision-Language Models”
- LinkedLinked via arxiv author · 85%Jeonghoon Mo →
“Context-structured Video Anomaly Detection with Large Vision-Language Models”
