Video Action Detection on JHMDB21 1.0 (test)
90.2f-mAP@0.5EVAD
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| EVADBackbone=ViT-B, Learning Supervision=Fully-Supervised, Annotation %=100%2024.12 | 90.2 | — | — | |
| BMVITBackbone=ViT-B, Learning Supervision=Fully-Supervised, Annotation %=100%2024.12 | 88.4 | — | — | |
| TubeRBackbone=I3D, Learning Supervision=Fully-Supervised, Annotation %=100%2024.12 | 87.4 | — | 82.3 | |
| STMixerBackbone=SF-R101NL, Learning Supervision=Fully-Supervised, Annotation %=100%2024.12 | 86.7 | — | — | |
| ACAR-NetBackbone=SF-R50, Learning Supervision=Fully-Supervised, Annotation %=100%2024.12 | 77.9 | — | 80.1 | |
| YOWOBackbone=ResNext-101, Learning Supervision=Fully-Supervised, Annotation %=100%2024.12 | 74.4 | 85.7 | 58.1 | |
| MOCBackbone=DLA-34, Learning Supervision=Fully-Supervised, Annotation %=100%2024.12 | 70.8 | 77.3 | 70.2 | |
| Stable Mean TeacherBackbone=I3D, Learning Supervision=Semi-Supervised, Annotation %=20%2024.12 | 69.8 | 98.8 | 70.7 | |
| GLNetBackbone=I3D, Learning Supervision=Weakly-Supervised2024.12 | 65.9 | 77.3 | 50.8 | |
| TACNetBackbone=RN-50, Learning Supervision=Fully-Supervised, Annotation %=100%2024.12 | 65.5 | 74.1 | 73.4 | |
| VideoCapsuleNetBackbone=I3D, Learning Supervision=Fully-Supervised, Annotation %=100%2024.12 | 64.6 | 95.1 | — | |
| E2E-SSLBackbone=I3D, Learning Supervision=Semi-Supervised, Annotation %=20%2024.12 | 59.1 | 93.2 | 58.7 | |
| ISDBackbone=I3D, Learning Supervision=Semi-Supervised, Annotation %=20%2024.12 | 57.8 | 90.2 | 57 | |
| Baseline Mean TeacherBackbone=I3D, Learning Supervision=Semi-Supervised, Annotation %=20%2024.12 | 56.3 | 88.8 | 52.8 | |
| Supervised baselineBackbone=I3D, Learning Supervision=Semi-Supervised, Annotation %=20%2024.12 | 55.7 | 93.9 | 52.4 | |
| Pseudo-labelBackbone=I3D, Learning Supervision=Semi-Supervised, Annotation %=20%2024.12 | 55.3 | 87.6 | 52 | |
| MixMatchBackbone=I3D, Learning Supervision=Semi-Supervised, Annotation %=30%2024.12 | 7.5 | 46.2 | 5.8 |