Violence Detection on RWF-2000 (test)
82.8AccuracyMobileNetV2+BiLSTM
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| MobileNetV2+BiLSTMTraining Scenario=E7 All-source2026.05 | 82.8 | 90.1 | |
| MobileNetV2+Temporal CNNTraining Scenario=E7 All-source2026.05 | 81.8 | 90.6 | |
| MobileNetV2+LSTMTraining Scenario=E7 All-source2026.05 | 81.1 | 90.8 | |
| Fuse Mamba-VDModel Type=SSM2025.05 | 0.945 | — | |
| FuseMamba-VD2025.05 | 0.945 | — | |
| Multi-Head Att & LSTMModel Type=LSTM + ViViT2025.05 | 0.9436 | — | |
| Reported Best2025.05 | 0.9436 | — | |
| CUE-NetModel Type=Enhanced UniformerV22025.05 | 0.94 | — | |
| ACTION-VSTModel Type=CNN + ViViT2025.05 | 0.9359 | — | |
| Structured Keypoint PoolingModel Type=CNN2025.05 | 0.934 | — | |
| VideoMambaModel Type=SSM2025.05 | 0.9275 | — | |
| Video Swin TransformerModel Type=ViViT2025.05 | 0.9125 | — | |
| SPILModel Type=Graph CNN2025.05 | 0.893 | — | |
| Flow Gated NetworkModel Type=Two Stream Graph CNN2025.05 | 0.8725 | — | |
| X3DModel Type=3DCNN2025.05 | 0.8475 | — | |
| I3DModel Type=3DCNN2025.05 | 0.834 | — | |
| ConvLSTMModel Type=CNN+LSTM2025.05 | 0.77 | — |