Video Anomaly Detection on UCF-Crime (test)
94.36AUCVANGUARD
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| VANGUARDTraining Stage=Stage-12026.04 | 94.36 | — | — | — | — | — | — | 77.5 | |
| VANGUARD2026.04 | 93.78 | — | — | — | — | — | — | 83.6 | |
| ASK-Hint2026.04 | 89.83 | — | — | — | — | — | — | — | |
| MemoVADBackbone=VideoMAE-S2026.06 | 89.45 | — | — | — | — | — | — | — | |
| Holmes-VAUBackbone=ViT, VAD Category=Explainable Multi-modal2024.12 | 88.96 | — | — | — | — | — | — | — | |
| Holmes-VAUExplainability=Explainable, Category=Fine-tuning2026.05 | 88.96 | — | — | — | — | — | — | — | |
| DAKDSFeatures=CLIP, T=32=true2024.06 | 88.34 | — | — | — | — | — | — | — | |
| Ex-VADVenue=ICML’25, Supervision=Weak, Explanation=✓, Training Data=100% / 100%2026.07 | 88.29 | — | — | — | — | — | — | — | |
| DAKDTFeatures=Multiple, T=32=true2024.06 | 88.15 | — | — | — | — | — | — | — | |
| DAKDSFeatures=I3D, T=32=true2024.06 | 88.1 | — | — | — | — | — | — | — | |
| VADorExplainability=Explainable, Category=Fine-tuning2026.05 | 88.1 | — | — | — | — | — | — | — | |
| VadCLIPBackbone=ViT, VAD Category=Non-explainable2024.12 | 88.02 | — | — | — | — | — | — | — | |
| VADCLIPExplainability=Non-explainable, Category=Weakly supervised2026.05 | 88.02 | — | — | — | — | — | — | — | |
| VadCLIPBackbone=CLIP2026.06 | 88.02 | — | — | — | — | — | — | — | |
| VadCLIPVenue=AAAI’24, Supervision=Weak, Explanation=X, Training Data=100% / 100%2026.07 | 88.02 | — | — | — | — | — | — | — | |
| TPWNGSupervision=Weakly2024.04 | 87.79 | — | — | — | — | — | — | — | |
| Yang et al.Backbone=ViT, VAD Category=Non-explainable2024.12 | 87.79 | — | — | — | — | — | — | — | |
| TPWNGBackbone=CLIP2026.06 | 87.79 | — | — | — | — | — | — | — | |
| CLIP-TSAMode=Weak, use_entire_training_set=true2023.11 | 87.58 | — | — | — | — | — | — | — | |
| CLIP-TSASupervision=Weakly2024.04 | 87.58 | — | — | — | — | — | — | — | |
| CLIP-TSABackbone=ViT, VAD Category=Non-explainable2024.12 | 87.58 | — | — | — | — | — | — | — | |
| CLIP-TSAExplainability=Non-explainable2024.12 | 87.58 | — | — | — | — | — | — | — | |
| CLIP-TSAExplainability=Non-explainable, Category=Weakly supervised2026.05 | 87.58 | — | — | — | — | — | — | — | |
| CLIP-TSABackbone=CLIP2026.06 | 87.58 | — | — | — | — | — | — | — | |
| EventVADBackbone=CLIP2026.06 | 87.51 | — | — | — | — | — | — | — | |
| SSRLExplainability=Non-explainable2024.12 | 87.43 | — | — | — | — | — | — | — | |
| Flashback (2025)2026.04 | 87.29 | — | — | — | — | — | — | — | |
| MGFNBackbone=I3D, VAD Category=Non-explainable2024.12 | 86.98 | — | — | — | — | — | — | — | |
| MGFNExplainability=Non-explainable2024.12 | 86.98 | — | — | — | — | — | — | — | |
| MGFNFeatures=I3D-RGB, T=32=true2024.06 | 86.98 | — | — | — | — | — | — | — | |
| MGFNExplainability=Non-explainable, Category=Weakly supervised2026.05 | 86.98 | — | — | — | — | — | — | — | |
| UR-DMUSupervision=Weakly2024.04 | 86.97 | — | — | — | — | — | — | — | |
| UR-DMUBackbone=I3D, VAD Category=Non-explainable2024.12 | 86.97 | — | — | — | — | — | — | — | |
| UR-DMUBackbone=I3D2026.06 | 86.97 | — | — | — | — | — | — | — | |
| SSRL*Features=I3D-RGB, T=32=true2024.06 | 86.79 | — | — | — | — | — | — | — | |
| PEL4VADVenue=TIP’24, Supervision=Weak, Explanation=X, Training Data=100% / 100%2026.07 | 86.76 | — | — | — | — | — | — | — | |
| DMUMode=Weak, use_entire_training_set=true2023.11 | 86.75 | — | — | — | — | — | — | — | |
| UMILMode=Weak, use_entire_training_set=true2023.11 | 86.75 | — | — | — | — | — | — | — | |
| MGFNSupervision=Weakly2024.04 | 86.67 | — | — | — | — | — | — | — | |
| MGFNBackbone=I3D2026.06 | 86.67 | — | — | — | — | — | — | — | |
| LRPOSupervision=Verbalized Learning, Explanation=✓, Training Data=100% / 100%2026.07 | 86.59 | — | — | — | — | — | — | — | |
| VERAExplainability=Explainable, Backbone=InternVL2-8B2024.12 | 86.55 | — | — | — | — | — | — | — | |
| VERA2026.04 | 86.55 | — | — | — | — | — | — | — | |
| VERAExplainability=Explainable, Category=Verbalized2026.05 | 86.55 | — | — | — | — | — | — | — | |
| VERAVenue=CVPR’25, Supervision=Verbalized Learning, Explanation=✓, Training Data=100% / 100%2026.07 | 86.55 | — | — | — | — | — | — | — | |
| RTFMLearning Protocol=MTL2025.12 | 86.5 | — | — | — | — | — | — | — | |
| OVVADMode=Weak, use_entire_training_set=false2023.11 | 86.4 | — | — | — | 93.8 | 88.2 | — | — | |
| Wu et al.Backbone=ViT, VAD Category=Non-explainable2024.12 | 86.4 | — | — | — | — | — | — | — | |
| Wu et al.Features=CLIP2025.03 | 86.4 | — | — | — | 93.8 | 88.2 | — | — | |
| OVVADExplainability=Non-explainable2024.12 | 86.4 | — | — | — | — | — | — | — | |
| OVVADBackbone=CLIP2026.06 | 86.4 | — | — | — | — | — | — | — | |
| SphereVADLLM/VLM Calls (per video)=0 (feature extraction only), GPU Inference (total GPU-hours)=0.56 h (4×A100)‡, Post-Extraction Time=4.5 s (CPU)2026.05 | 86.38 | — | — | — | — | — | — | — | |
| Zhang et al.Supervision=Weakly2024.04 | 86.22 | — | — | — | — | — | — | — | |
| S3RBackbone=I3D, VAD Category=Non-explainable2024.12 | 85.99 | — | — | — | — | — | — | — | |
| S3RExplainability=Non-explainable2024.12 | 85.99 | — | — | — | — | — | — | — | |
| S3RFeatures=I3D-RGB, T=32=true2024.06 | 85.99 | — | — | — | — | — | — | — | |
| S3RExplainability=Non-explainable, Category=Weakly supervised2026.05 | 85.99 | — | — | — | — | — | — | — | |
| VADorExplainability=Explainable, Instruction Tuning=false2024.12 | 85.9 | — | — | — | — | — | — | — | |
| RTFMMode=Weak, use_entire_training_set=true2023.11 | 85.66 | — | — | — | — | — | — | — | |
| MSLSupervision=Weakly2024.04 | 85.62 | — | — | — | — | — | — | — | |
| MSLExplainability=Non-explainable2024.12 | 85.62 | — | — | — | — | — | — | — | |
| MSLFeatures=VSwin-RGB, T=32=true2024.06 | 85.62 | — | — | — | — | — | — | — | |
| MSLExplainability=Non-explainable, Category=Weakly supervised2026.05 | 85.62 | — | — | — | — | — | — | — | |
| MSLBackbone=I3D2026.06 | 85.62 | — | — | — | — | — | — | — | |
| LRPOSupervision=Verbalized Learning, Explanation=✓, Training Data=2.5% / 6%2026.07 | 85.36 | — | — | — | — | — | — | — | |
| MSLBackbone=I3D, VAD Category=Non-explainable2024.12 | 85.3 | — | — | — | — | — | — | — | |
| MSLFeatures=I3D-RGB, T=32=true2024.06 | 85.3 | — | — | — | — | — | — | — | |
| DMUMode=Weak, use_entire_training_set=false2023.11 | 85.14 | — | — | — | 93.52 | 86.24 | — | — | |
| CRFDSupervision=Weakly2024.04 | 84.89 | — | — | — | — | — | — | — | |
| Wu et al. [33]Features=I3D-RGB, T=32=true2024.06 | 84.89 | — | — | — | — | — | — | — | |
| PANDALLM/VLM Calls (per video)=Multiple (VLM + MLLM, iterative reflection) Average speed: 0.82 FPS on A60002026.05 | 84.89 | — | — | — | — | — | — | — | |
| CRFDBackbone=I3D2026.06 | 84.89 | — | — | — | — | — | — | — | |
| VADTreeLLM/VLM Calls (per video)=VLM caption + LLM scoring per node, GPU Inference (total GPU-hours)=62.3 h (2×3090)2026.05 | 84.74 | — | — | — | — | — | — | — | |
| VADTreeVenue=NeurIPS’25, Supervision=Training-free, Explanation=✓, Training Data=-2026.07 | 84.74 | — | — | — | — | — | — | — | |
| Holmes-VADExplainability=Explainable, Instruction Tuning=false2024.12 | 84.61 | — | — | — | — | — | — | — | |
| Wu et al.Mode=Weak, use_entire_training_set=true2023.11 | 84.57 | — | — | — | — | — | — | — | |
| DYANNETBackbone=I3D, VAD Category=Non-explainable2024.12 | 84.5 | — | — | — | — | — | — | — | |
| DYANNETExplainability=Non-explainable2024.12 | 84.5 | — | — | — | — | — | — | — | |
| AnomizeFeatures=CLIP2025.03 | 84.49 | — | — | — | 93 | 87.05 | — | — | |
| RTFMMode=Weak, use_entire_training_set=false2023.11 | 84.47 | — | — | — | 92.54 | 85.87 | — | — | |
| RTFMFeatures=CLIP2025.03 | 84.47 | — | — | — | 92.54 | 85.87 | — | — | |
| VADR12026.04 | 84.45 | — | — | — | — | — | — | 85.06 | |
| RTFMSupervision=Weakly2024.04 | 84.3 | — | — | — | — | — | — | — | |
| RTFMBackbone=I3D, VAD Category=Non-explainable2024.12 | 84.3 | — | — | — | — | — | — | — | |
| RTFMExplainability=Non-explainable2024.12 | 84.3 | — | — | — | — | — | — | — | |
| RTFM*Features=I3D-RGB, T=32=true2024.06 | 84.3 | — | — | — | — | — | — | — | |
| RTFMExplainability=Non-explainable, Category=Weakly supervised2026.05 | 84.3 | — | — | — | — | — | — | — | |
| RTFMBackbone=I3D2026.06 | 84.3 | — | — | — | — | — | — | — | |
| RFTMVenue=ICCV’21, Supervision=Weak, Explanation=X, Training Data=100% / 100%2026.07 | 84.3 | — | — | — | — | — | — | — | |
| Unified-VADLLM/VLM Calls (per video)=1 VLM + 1–2 LLM per 16 frames Amortised: 0.029 s/frame on 2×30902026.05 | 84.28 | — | — | — | — | — | — | — | |
| URFVenue=NeurIPS’25, Supervision=Training-free, Explanation=✓, Training Data=-2026.07 | 84.28 | — | — | — | — | — | — | — | |
| Sultani et al.Mode=Weak, use_entire_training_set=true2023.11 | 84.14 | — | — | — | — | — | — | — | |
| SpatialVLM2026.04 | 83.18 | — | — | — | — | — | — | 47.83 | |
| CLAWSExplainability=Non-explainable2024.12 | 83.03 | — | — | — | — | — | — | — | |
| CLAWSFeatures=C3D - RGB, T=32=false2024.06 | 83.03 | — | — | — | — | — | — | — | |
| CoReVADExplainability=Explainable, Category=Training-free2026.05 | 82.51 | — | — | — | — | — | — | — | |
| MCANetExplainability=Explainable, Category=Training-free2026.05 | 82.47 | — | — | — | — | — | — | — | |
| AVVDMode=Weak, use_entire_training_set=true2023.11 | 82.45 | — | — | — | — | — | — | — | |
| HL-NetSupervision=Weakly2024.04 | 82.44 | — | — | — | — | — | — | — | |
| Wu et al.Backbone=I3D, VAD Category=Non-explainable2024.12 | 82.44 | — | — | — | — | — | — | — |