Action Recognition on Kinetics (test)
96.4Average AccOtter
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| OtterNumber of shots=5-shot, Evaluation protocol=Regular testing, Training source dataset=Kinetics2025.11 | 96.4 | — | — | — | |
| DISTShot=5-shot, Backbone=CLIP ViT-B, N-way=5-way, Evaluation Protocol=combined few-shot and zero-shot2026.02 | 96 | — | — | — | |
| DISTShot=1-shot, Backbone=CLIP ViT-B, N-way=5-way, Evaluation Protocol=combined few-shot and zero-shot2026.02 | 95.6 | — | — | — | |
| CLIP-FSARShot=5-shot, Backbone=CLIP ViT-B, N-way=5-way, Evaluation Protocol=combined few-shot and zero-shot2026.02 | 95.4 | — | — | — | |
| CLIP-FSARShot=1-shot, Backbone=CLIP ViT-B, N-way=5-way, Evaluation Protocol=combined few-shot and zero-shot2026.02 | 94.8 | — | — | — | |
| MantaNumber of shots=5-shot, Evaluation protocol=Regular testing, Training source dataset=Kinetics2025.11 | 94.2 | — | — | — | |
| OtterNumber of shots=5-shot, Evaluation protocol=Cross-dataset testing, Training source dataset=SSv22025.11 | 94 | — | — | — | |
| SOAPNumber of shots=5-shot, Evaluation protocol=Regular testing, Training source dataset=Kinetics2025.11 | 93.8 | — | — | — | |
| MantaNumber of shots=5-shot, Evaluation protocol=Cross-dataset testing, Training source dataset=SSv22025.11 | 91.8 | — | — | — | |
| SOAPNumber of shots=5-shot, Evaluation protocol=Cross-dataset testing, Training source dataset=SSv22025.11 | 91.1 | — | — | — | |
| OtterNumber of shots=1-shot, Evaluation protocol=Regular testing, Training source dataset=Kinetics2025.11 | 90.5 | — | — | — | |
| OtterNumber of shots=1-shot, Evaluation protocol=Cross-dataset testing, Training source dataset=SSv22025.11 | 89.2 | — | — | — | |
| MantaNumber of shots=1-shot, Evaluation protocol=Regular testing, Training source dataset=Kinetics2025.11 | 87.4 | — | — | — | |
| MantaNumber of shots=1-shot, Evaluation protocol=Cross-dataset testing, Training source dataset=SSv22025.11 | 86.3 | — | — | — | |
| SOAPNumber of shots=1-shot, Evaluation protocol=Regular testing, Training source dataset=Kinetics2025.11 | 86.1 | — | — | — | |
| MoLoNumber of shots=5-shot, Evaluation protocol=Regular testing, Training source dataset=Kinetics2025.11 | 85.7 | — | — | — | |
| SOAPNumber of shots=1-shot, Evaluation protocol=Cross-dataset testing, Training source dataset=SSv22025.11 | 84.1 | — | — | — | |
| NL I3Dbackbone=ResNet-101, modality=RGB2017.11 | 83.8 | — | — | — | |
| MoLoNumber of shots=5-shot, Evaluation protocol=Cross-dataset testing, Training source dataset=SSv22025.11 | 83.2 | — | — | — | |
| 2-Stream I3Dbackbone=Inception, modality=RGB + flow2017.11 | 82.8 | 74.2 | 91.3 | — | |
| Two-stream I3DInput modality=Two-stream, Training protocol=from scratch2017.11 | 80.8 | 71.6 | 90 | — | |
| I3Dbackbone=Inception, modality=RGB2017.11 | 80.2 | 71.1 | 89.3 | — | |
| ResNeXt-101 (64f)Input resolution/frames=64x112x112, Training protocol=from scratch2017.11 | 78.4 | — | — | — | |
| RGB-I3DInput modality=RGB, Input resolution/frames=64x224x224, Training protocol=from scratch2017.11 | 78.2 | 68.4 | 88 | — | |
| ResNeXt-101Input resolution/frames=16x112x112, Training protocol=from scratch2017.11 | 74.5 | — | — | — | |
| MoLoNumber of shots=1-shot, Evaluation protocol=Regular testing, Training source dataset=Kinetics2025.11 | 74.2 | — | — | — | |
| MoLoNumber of shots=1-shot, Evaluation protocol=Cross-dataset testing, Training source dataset=SSv22025.11 | 71.5 | — | — | — | |
| Two-stream CNNInput modality=Two-stream2017.11 | 71.2 | 61 | 81.3 | — | |
| CNN+LSTM2017.11 | 68 | 57 | 79 | — | |
| C3D w/ BNBatch Normalization (BN)=true2017.11 | 67.8 | 56.1 | 79.5 | — | |
| 2s-AGCN2020.10 | — | 36.1 | 58.7 | — | |
| A2-Net#Frames=8, FLOPs=40.8 G2018.10 | — | 74.6 | 91.5 | — | |
| ARTNet w/o TSNSpatial resolution=112 × 112, Backbone architecture=ResNet-182017.11 | — | — | — | 77.3 | |
| ARTNet with TSNSpatial resolution=112 × 112, Backbone architecture=ResNet-182017.11 | — | — | — | 78.7 | |
| AS-GCN2019.04 | — | 34.8 | 56.5 | — | |
| AS-GCN2020.10 | — | 34.8 | 56.5 | — | |
| C3DSpatial resolution=112 × 112, Backbone architecture=VGGNet-112017.11 | — | — | — | 67.8 | |
| C3DSpatial resolution=112 × 112, Backbone architecture=ResNet-182017.11 | — | — | — | 74.4 | |
| C3DSpatial resolution=112 × 112, Backbone architecture=ResNet-342017.11 | — | — | — | 75.3 | |
| Conv2020.10 | — | 30.8 | 52.6 | — | |
| Conv-Chiral2020.10 | — | 30.9 | 53 | — | |
| ConvNet+LSTMSpatial resolution=299 × 299, Backbone architecture=ResNet-502017.11 | — | — | — | 68 | |
| ConvNet+LSTM2018.10 | — | 63.3 | — | — | |
| Deep LSTM2019.04 | — | 16.4 | 35.3 | — | |
| Deep LSTM2020.10 | — | 16.4 | 35.3 | — | |
| Feature Enc2019.04 | — | 14.9 | 25.8 | — | |
| Feature Enc2020.10 | — | 14.9 | 25.8 | — | |
| I3D#Frames=64, FLOPs=107.9 G2018.10 | — | 71.1 | 89.3 | — | |
| PR-GCN2020.10 | — | 33.7 | 55.8 | — | |
| R(2+1)D#Frames=32, FLOPs=152.4 G2018.10 | — | 72 | 90 | — | |
| RGB-I3DSpatial resolution=224 × 224, Backbone architecture=Inception V12017.11 | — | — | — | 78.2 | |
| ST-GCN2019.04 | — | 30.7 | 52.8 | — | |
| ST-GCN2020.10 | — | 30.7 | 52.8 | — | |
| TCN2020.10 | — | 20.3 | 40 | — | |
| Temporal Conv2019.04 | — | 20.3 | 40 | — | |
| Two Stream Spatial NetworksSpatial resolution=299 × 299, Backbone architecture=ResNet-502017.11 | — | — | — | 66.6 |