Action Recognition on HMDB51 (test)
37.6AccuracyTF-VAEGAN
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| TF-VAEGANProtocol=GZSL2020.03 | 37.6 | — | — | — | |
| CEWGANProtocol=GZSL2020.03 | 36.1 | — | — | — | |
| f-VAEGANProtocol=GZSL2020.03 | 35.6 | — | — | — | |
| TF-VAEGANProtocol=ZSL2020.03 | 33 | — | — | — | |
| CLSWGANProtocol=GZSL2020.03 | 32.7 | — | — | — | |
| f-VAEGANProtocol=ZSL2020.03 | 31.1 | — | — | — | |
| CEWGANProtocol=ZSL2020.03 | 30.2 | — | — | — | |
| CLSWGANProtocol=ZSL2020.03 | 29.1 | — | — | — | |
| Obj2ActProtocol=ZSL2020.03 | 24.5 | — | — | — | |
| GGMProtocol=ZSL2020.03 | 20.7 | — | — | — | |
| GGMProtocol=GZSL2020.03 | 20.1 | — | — | — | |
| DISTShot=5-shot, Backbone=CLIP ViT-B, N-way=5-way, Evaluation Protocol=combined few-shot and zero-shot2026.02 | 0.887 | — | — | — | |
| CLIP-FSARShot=5-shot, Backbone=CLIP ViT-B, N-way=5-way, Evaluation Protocol=combined few-shot and zero-shot2026.02 | 0.877 | — | — | — | |
| PoseConv3DModality=RGB + Flow + Pose, Fusion=LateFusion, Pre-training=Kinetics4002021.04 | 0.85 | — | — | — | |
| SMARTBackbone=TSN+Kinetics2020.12 | 0.843 | — | — | — | |
| DynaMotion + I3DBackbone=Inc v32020.12 | 0.842 | — | — | — | |
| OmniSourceModality=RGB + Flow2021.04 | 0.838 | — | — | — | |
| DISTShot=1-shot, Backbone=CLIP ViT-B, N-way=5-way, Evaluation Protocol=combined few-shot and zero-shot2026.02 | 0.826 | — | — | — | |
| I3DBackbone=Inc v32020.12 | 0.807 | — | — | — | |
| KI-NetBackbone=Res-1522020.12 | 0.782 | — | — | — | |
| MMEBackbone=ViT-B, Pre-train=K4002022.10 | 0.78 | — | — | — | |
| AASBackbone=TSN+Kinetics2020.12 | 0.773 | — | — | — | |
| CLIP-FSARShot=1-shot, Backbone=CLIP ViT-B, N-way=5-way, Evaluation Protocol=combined few-shot and zero-shot2026.02 | 0.771 | — | — | — | |
| Fully supervisedArchitecture=S3D-G, Pre-train Dataset=Kinetics-4002020.10 | 0.759 | — | — | — | |
| Prob-Distill+Flow I3D*Input frames=24, Backbone=I3D2019.04 | 0.757 | — | — | — | |
| DUALPATHBackbone=ViT-B/16, Classifier=MLPs, Learnable Parameters=10M2023.03 | 0.756 | — | — | — | |
| I3D-RGBFLOPs=107.9 G, +OF=false2018.07 | 0.748 | — | — | — | |
| I3D-RGB*Pretraining Dataset=Kinetics + Imagenet, Pretraining Supervision=fully supervised (object+action labels), Source=as listed in [18]2018.06 | 0.748 | — | — | — | |
| Two Stream I3D*Input frames=24, Backbone=I3D, Modality=Two Stream2019.04 | 0.748 | — | — | — | |
| MF-NetFLOPs=11.1 G, +OF=false2018.07 | 0.746 | — | — | — | |
| SMARTBackbone=TSN2020.12 | 0.746 | — | — | — | |
| R(2+1)D-RGBFLOPs=152.4 G, +OF=false2018.07 | 0.745 | — | — | — | |
| MARS+Flow ResNeXtBackbone=ResNeXt2019.04 | 0.745 | — | — | — | |
| I3D-RGB*Pretraining Dataset=Kinetics, Pretraining Supervision=fully supervised (action labels), Source=as listed in [18]2018.06 | 0.743 | — | — | — | |
| (2+1)D ResNet-50Supervision=Supervised, Pre-training=Kinetics2020.02 | 0.743 | — | — | — | |
| Two Stream ResNeXtBackbone=ResNeXt, Modality=Two Stream2019.04 | 0.74 | — | — | — | |
| VideoMAEBackbone=ViT-B, Pre-train=K4002022.10 | 0.733 | — | — | — | |
| TC-CLIPK (shots)=16, Pre-trained=Kinetics-400, Fine-tuning=true2024.04 | 0.73 | — | — | — | |
| GDTBackbone=R(2+1)D, Pre-train=IG65M2022.10 | 0.728 | — | — | — | |
| MARS2019.04 | 0.723 | — | — | — | |
| ρBYOLBackbone=R50, Pre-train=K4002022.10 | 0.721 | — | — | — | |
| Prob-DistillBackbone=RGB backbone2019.04 | 0.72 | — | — | — | |
| VL PromptingK (shots)=16, Pre-trained=Kinetics-4002022.12 | 0.72 | — | — | — | |
| ViFi-CLIPK (shots)=16, Pre-trained=Kinetics-400, Fine-tuning=true2024.04 | 0.72 | — | — | — | |
| XCLIPK (shots)=16, Pre-trained=Kinetics-4002022.12 | 0.717 | — | — | — | |
| X-CLIPK (shots)=16, Pre-trained=Kinetics-400, Fine-tuning=true2024.04 | 0.717 | — | — | — | |
| TC-CLIPK (shots)=8, Pre-trained=Kinetics-400, Fine-tuning=true2024.04 | 0.714 | — | — | — | |
| AASBackbone=TSN2020.12 | 0.712 | — | — | — | |
| TSNNumber of segments=72017.05 | 0.71 | — | — | — | |
| TVNet2019.04 | 0.71 | — | — | — | |
| ARTNetFLOPs=25.7 G, +OF=false2018.07 | 0.709 | — | — | — | |
| I3D RGB*Input frames=24, Backbone=I3D, Modality=RGB2019.04 | 0.709 | — | — | — | |
| TSNNumber of segments=32017.05 | 0.707 | — | — | — | |
| FeatMatch2019.04 | 0.707 | — | — | — | |
| TSNBackbone=BN-Inc2020.12 | 0.699 | — | — | — | |
| VL PromptingK (shots)=8, Pre-trained=Kinetics-4002022.12 | 0.696 | — | — | — | |
| ViFi-CLIPK (shots)=8, Pre-trained=Kinetics-400, Fine-tuning=true2024.04 | 0.696 | — | — | — | |
| TSNFLOPs=3.8 G, +OF=true2018.07 | 0.694 | — | — | — | |
| PoseConv3DModality=Pose, Pre-training=Kinetics4002021.04 | 0.693 | — | — | — | |
| XCLIPK (shots)=8, Pre-trained=Kinetics-4002022.12 | 0.693 | — | — | — | |
| X-CLIPK (shots)=8, Pre-trained=Kinetics-400, Fine-tuning=true2024.04 | 0.693 | — | — | — | |
| TC-CLIPK (shots)=4, Pre-trained=Kinetics-400, Fine-tuning=true2024.04 | 0.685 | — | — | — | |
| CORPBackbone=R50, Pre-train=K4002022.10 | 0.68 | — | — | — | |
| CVRLBackbone=R50, Pre-train=K4002022.10 | 0.679 | — | — | — | |
| Evolved LossSupervision=Weakly guided2020.02 | 0.678 | — | — | — | |
| ELoSupervision=Unsupervised2020.02 | 0.674 | — | — | — | |
| XDCBackbone=R(2+1)D, Pre-train=IG65M2022.10 | 0.671 | — | — | — | |
| MC3Pretraining Dataset=Kinetics, Pretraining Supervision=fully supervised (action labels)2018.06 | 0.668 | — | — | — | |
| XCLIPK (shots)=4, Pre-trained=Kinetics-4002022.12 | 0.668 | — | — | — | |
| X-CLIPK (shots)=4, Pre-trained=Kinetics-400, Fine-tuning=true2024.04 | 0.668 | — | — | — | |
| VideoPromptBackbone=ViT-B/16, Classifier=Temporal transformer, Learnable Parameters=6M2023.03 | 0.664 | — | — | — | |
| ActionCLIPK (shots)=16, Pre-trained=Kinetics-4002022.12 | 0.661 | — | — | — | |
| ActionCLIPK (shots)=16, Pre-trained=Kinetics-400, Fine-tuning=true2024.04 | 0.661 | — | — | — | |
| ST-AdapterBackbone=ViT-B/16, Classifier=Linear, Learnable Parameters=7M2023.03 | 0.659 | — | — | — | |
| A5K (shots)=16, Pre-trained=Kinetics-4002022.12 | 0.658 | — | — | — | |
| A5K (shots)=16, Pre-trained=Kinetics-400, Fine-tuning=true2024.04 | 0.658 | — | — | — | |
| MPR2017.05 | 0.655 | — | — | — | |
| TC-CLIPK (shots)=2, Pre-trained=Kinetics-400, Fine-tuning=true2024.04 | 0.653 | — | — | — | |
| VL PromptingK (shots)=4, Pre-trained=Kinetics-4002022.12 | 0.651 | — | — | — | |
| ViFi-CLIPK (shots)=4, Pre-trained=Kinetics-400, Fine-tuning=true2024.04 | 0.651 | — | — | — | |
| LTC2017.05 | 0.648 | — | — | — | |
| RSPNetArchitecture=S3D-G, Pre-train Dataset=Kinetics-400, Pre-train Epochs=10002020.10 | 0.647 | — | — | — | |
| VideoDarwin2017.05 | 0.637 | — | — | — | |
| AdaptFormerBackbone=ViT-B/16, Classifier=Temporal transformer, Learnable Parameters=8M2023.03 | 0.637 | — | — | — | |
| KVMF2017.05 | 0.633 | — | — | — | |
| Pro-tuningBackbone=ViT-B/16, Classifier=Temporal transformer, Learnable Parameters=9M2023.03 | 0.633 | — | — | — | |
| TDD+FV2017.05 | 0.632 | — | — | — | |
| VL PromptingK (shots)=2, Pre-trained=Kinetics-4002022.12 | 0.63 | — | — | — | |
| ViFi-CLIPK (shots)=2, Pre-trained=Kinetics-400, Fine-tuning=true2024.04 | 0.63 | — | — | — | |
| Two-streamBackbone=VGG2020.12 | 0.624 | — | — | — | |
| VPTBackbone=ViT-B/16, Classifier=Temporal transformer, Learnable Parameters=7M2023.03 | 0.624 | — | — | — | |
| MC2Pretraining Dataset=Kinetics, Pretraining Supervision=fully supervised (action labels)2018.06 | 0.62 | — | — | — | |
| MoFAP2017.05 | 0.617 | — | — | — | |
| MC3Pretraining Dataset=Audioset, Pretraining Supervision=self-supervised (AVTS)2018.06 | 0.616 | — | — | — | |
| AVTSSupervision=Unsupervised2020.02 | 0.616 | — | — | — | |
| Dynamic Image2019.04 | 0.613 | — | — | — | |
| A5K (shots)=8, Pre-trained=Kinetics-4002022.12 | 0.613 | — | — | — | |
| A5K (shots)=8, Pre-trained=Kinetics-400, Fine-tuning=true2024.04 | 0.613 | — | — | — | |
| Linear w/ ViT-B/16Backbone=ViT-B/16, Classifier=Linear, Learnable Parameters=0.1M2023.03 | 0.612 | — | — | — | |
| iDT+HSV2017.05 | 0.611 | — | — | — |