Violence Detection on XD-Violence
87.61APDSRL
Evaluation Results
| Method | Links | |
|---|---|---|
| DSRLInput Setting=Multimodal, Feature Space=Euclidean and Hyperbolic2024.09 | 87.61 | |
| HyperVDInput Setting=Multimodal, Feature Space=Hyperbolic2024.09 | 85.67 | |
| MSBTModality=RGB+Audio+Flow2024.05 | 84.32 | |
| Ours (full)Learning Manner=Weakly-supervised, Modality=A+V, Parameters=0.678M2022.07 | 83.4 | |
| Yu et al. (Full)Modality=RGB+Audio2024.05 | 83.4 | |
| MACIL-SDInput Setting=Multimodal, Feature Space=Euclidean2024.09 | 83.4 | |
| Xiao et al.Modality=RGB+Audio+Flow2024.05 | 83.09 | |
| MSBTModality=RGB+Audio2024.05 | 82.54 | |
| Ours (light)Learning Manner=Weakly-supervised, Modality=A+V, Parameters=0.347M2022.07 | 82.17 | |
| Yu et al. (Light)Modality=RGB+Audio2024.05 | 82.17 | |
| DSRLInput Setting=Unimodal, Feature Space=Euclidean and Hyperbolic2024.09 | 82.01 | |
| UMILInput Setting=Multimodal, Feature Space=Euclidean2024.09 | 81.77 | |
| Pang et al.Learning Manner=Weakly-supervised, Modality=A+V, Parameters=1.876M2022.07 | 81.69 | |
| UMILInput Setting=Unimodal, Feature Space=Euclidean2024.09 | 81.66 | |
| Completeness-and-Uncertainty Aware Pseudo Label EnhancementSupervision=Weakly, Feature=I3D+VGGish2022.12 | 81.43 | |
| Zhang et al.Modality=RGB+Audio2024.05 | 81.43 | |
| Zhang et al.Input Setting=Multimodal, Feature Space=Euclidean2024.09 | 81.43 | |
| MSBTModality=RGB+Flow2024.05 | 80.68 | |
| ACFModality=RGB+Audio2024.05 | 80.13 | |
| Yu et al. (Full)Modality=RGB+Flow, re-trained=true2024.05 | 79.73 | |
| Wu et al.Modality=RGB+Audio+Flow2024.05 | 79.53 | |
| Pang et al.Modality=RGB+Audio2024.05 | 79.37 | |
| Pang et al.Input Setting=Multimodal, Feature Space=Euclidean2024.09 | 79.37 | |
| MGFNInput Setting=Unimodal, Feature Space=Euclidean2024.09 | 79.19 | |
| Completeness-and-Uncertainty Aware Pseudo Label EnhancementSupervision=Weakly, Feature=I3D2022.12 | 78.74 | |
| CU-NetInput Setting=Unimodal, Feature Space=Euclidean2024.09 | 78.74 | |
| Wu et al.† [62]Learning Manner=Weakly-supervised, Modality=A+V, Parameters=1.539M, Re-implementation=integrating logits of two identical networks with audio and visual inputs2022.07 | 78.66 | |
| Wu et al. [62]Learning Manner=Weakly-supervised, Modality=A+V, Parameters=0.843M2022.07 | 78.64 | |
| Wu et al.Supervision=Weakly, Feature=I3D+VGGish2022.12 | 78.64 | |
| HL-NetModality=RGB+Audio2024.05 | 78.64 | |
| HL-NetInput Setting=Multimodal, Feature Space=Euclidean2024.09 | 78.64 | |
| Wu et al. [29]Input Setting=Multimodal, Feature Space=Euclidean2024.09 | 78.64 | |
| RTFM†Learning Manner=Weakly-supervised, Modality=A+V, Parameters=13.190M, Re-implementation=integrating logits of two identical networks with audio and visual inputs2022.07 | 78.54 | |
| Yu et al. (Light)Modality=RGB+Flow, re-trained=true2024.05 | 78.49 | |
| Li et al.Learning Manner=Weakly-supervised, Modality=V2022.07 | 78.28 | |
| MSLSupervision=Weakly, Feature=I3D2022.12 | 78.28 | |
| MSLInput Setting=Unimodal, Feature Space=Euclidean2024.09 | 78.28 | |
| RTFM*Learning Manner=Weakly-supervised, Modality=A+V, Parameters=13.510M, Re-implementation=fusing audio and visual features as inputs2022.07 | 78.1 | |
| RTFMLearning Manner=Weakly-supervised, Modality=V, Parameters=12.067M2022.07 | 77.81 | |
| RTFMSupervision=Weakly, Feature=I3D2022.12 | 77.81 | |
| RTFMInput Setting=Unimodal, Feature Space=Euclidean2024.09 | 77.81 | |
| MSBTModality=Audio+Flow2024.05 | 77.47 | |
| Yu et al. (Full)Modality=Audio+Flow, re-trained=true2024.05 | 77.06 | |
| Yu et al. (Light)Modality=Audio+Flow, re-trained=true2024.05 | 76.58 | |
| Wu et al. [61]Learning Manner=Weakly-supervised, Modality=V2022.07 | 75.9 | |
| Wu et al.Supervision=Weakly, Feature=I3D+VGGish2022.12 | 75.9 | |
| Wu et al. [27]Input Setting=Unimodal, Feature Space=Euclidean2024.09 | 75.9 | |
| Wu et al.Supervision=Weakly, Feature=I3D2022.12 | 75.41 | |
| Sultani et al.Learning Manner=Weakly-supervised, Modality=V2022.07 | 73.2 | |
| Sultani et al.Supervision=Weakly, Feature=I3D+VGGish2022.12 | 73.2 | |
| Sultani et al.Input Setting=Unimodal, Feature Space=Euclidean2024.09 | 73.2 | |
| Wu et al.Modality=Audio+Flow2024.05 | 72.96 | |
| SVM baselineLearning Manner=Unsupervised, Modality=V2022.07 | 50.78 | |
| SVM baselineSupervision=Semi, Feature=I3D+VGGish2022.12 | 50.78 | |
| Hasan et al.Learning Manner=Unsupervised, Modality=V2022.07 | 30.77 | |
| Hasan et al.Supervision=Semi, Feature=I3D+VGGish2022.12 | 30.77 | |
| OCSVMLearning Manner=Unsupervised, Modality=V2022.07 | 27.25 | |
| OCSVMSupervision=Semi, Feature=I3D+VGGish2022.12 | 27.25 |