Action Recognition on HMDB51 (Accuracy %)
93.25Accuracy (HMDB51)MSQNet
Evaluation Results
| Method | Links | |
|---|---|---|
| MSQNetBackbone=TS, Pretrain=K400, MM=Yes2023.07 | 93.25 | |
| EfficientNetB0ConvLSTMYear=20242025.01 | 89.22 | |
| Act-Control VisionYear=20242025.01 | 88.7 | |
| VideoMAE V2Modality=V2023.03 | 88.1 | |
| VideoMAE V2-gBackbone=ViT, Pretrain=K400/K600, MM=No2023.07 | 88.1 | |
| Attention-LSTM-3DCNNYear=20232025.01 | 87.98 | |
| OmniSourceInput=RGB/Flow2023.03 | 87 | |
| R2+1D-BERTBackbone=R(2+1)D, Pretrain=IG65M, MM=No2023.07 | 85.1 | |
| BIKEBackbone=ViT, Pretrain=WIT-400M, MM=Yes2023.07 | 84.31 | |
| STM FrameworkYear=20242025.01 | 80.4 | |
| 2-Attention CNN-GRUYear=20232025.01 | 79.38 | |
| Dual-Stream FrameworkYear=2024, Model Composition=VIT+PBILSTM+DSMHA2025.01 | 78.62 | |
| D Bi-LSTM Transfer LearningYear=20242025.01 | 76.3 | |
| Temporal Gradient LearningYear=20222025.01 | 75.96 | |
| DB-LSTMYear=20212025.01 | 75.17 | |
| TS-LSTMYear=20192025.01 | 74.86 | |
| I3DInput=RGB/Flow2023.03 | 74.8 | |
| ViTYear=20212025.01 | 73.75 | |
| MAE-Align w/ real humansPrivacy Preserving=false, Stage 1: MAE=Kinetics, Stage 2: Alignment=Kinetics, Evaluation Protocol=Fine-tuning, Backbone=ViT-B2023.11 | 73.4 | |
| VideoMAE V1Modality=V2023.03 | 73.3 | |
| GDTModality=V+A2023.03 | 72.8 | |
| Temporal Optical Flow LSTMYear=20192025.01 | 72.25 | |
| PBYOL_p=4Modality=V2023.03 | 72.1 | |
| Privacy-Preserving MAE-Align (PPMA)Privacy Preserving=true, Stage 1: MAE=NH Kinetics, Stage 2: Alignment=NH Kinetics + Synthetic, Evaluation Protocol=Fine-tuning, Backbone=ViT-B2023.11 | 71.2 | |
| OursInput=Skeleton2023.03 | 70.9 | |
| Temporal Segment NetworksYear=20192025.01 | 70.77 | |
| CVRLModality=V2023.03 | 70.6 | |
| Two-Stream 3DCNNYear=20182025.01 | 70.53 | |
| Deep Autoencoder CNNYear=20192025.01 | 70.37 | |
| PoseConv3DInput=Skeleton2023.03 | 69.7 | |
| MAE-Align w/ SyntheticPrivacy Preserving=true, Stage 1: MAE=Synthetic, Stage 2: Alignment=Synthetic, Evaluation Protocol=Fine-tuning, Backbone=ViT-B2023.11 | 69.7 | |
| MMVModality=V+A+T2023.03 | 69.6 | |
| MAE-Align w/ real humansPrivacy Preserving=false, Stage 1: MAE=Kinetics, Stage 2: Alignment=Kinetics, Evaluation Protocol=Linear Probing, Backbone=ViT-B2023.11 | 69.5 | |
| ViTYear=2023, config_variant=22025.01 | 68.2 | |
| CORP_fModality=V2023.03 | 68 | |
| ELOModality=V+A2023.03 | 67.4 | |
| Self-Supervised Transf (SVT)Year=2022, Evaluation Protocol=Fine-tune2025.01 | 67.28 | |
| XDCModality=V+A2023.03 | 67.1 | |
| Correlational ConvLSTMYear=20202025.01 | 66.28 | |
| SlowOnlyInput=RGB/Flow2023.03 | 66 | |
| Privacy-Preserving MAE-Align (PPMA)Privacy Preserving=true, Stage 1: MAE=NH Kinetics, Stage 2: Alignment=NH Kinetics + Synthetic, Evaluation Protocol=Linear Probing, Backbone=ViT-B2023.11 | 64.9 | |
| RSPNetModality=V2023.03 | 64.7 | |
| EquiAV2024.03 | 64.4 | |
| CPDModality=V+T2023.03 | 63.8 | |
| BraVeHigh temporal resolution=true, Industry-level computational resources=true2024.03 | 63.6 | |
| XKD2024.03 | 62.2 | |
| MIL-NCEModality=V+T2023.03 | 61 | |
| MMV2024.03 | 60 | |
| ViTYear=2023, config_variant=12025.01 | 59.74 | |
| MAE-Align w/ SyntheticPrivacy Preserving=true, Stage 1: MAE=Synthetic, Stage 2: Alignment=Synthetic, Evaluation Protocol=Linear Probing, Backbone=ViT-B2023.11 | 59.7 | |
| TimeSformer KineticsPrivacy Preserving=false, Stage 2: Alignment=Kinetics, Evaluation Protocol=Fine-tuning, Backbone=ViT-B, Pre-trained on ImageNet-21K=true2023.11 | 59.5 | |
| Self-Supervised Transf (SVT)Year=2022, Evaluation Protocol=Linear2025.01 | 57.8 | |
| InternVideo2_clip-6B#F=8, Training Data=IV-400M, protocol=Zero-shot2024.03 | 56.7 | |
| Criss Cross2024.03 | 56.2 | |
| XDC2024.03 | 56 | |
| Vi²CLRModality=V2023.03 | 55.7 | |
| I3D (RN50 backbone)Privacy Preserving=true, Stage 2: Alignment=Synthetic, Evaluation Protocol=Fine-tuning, Backbone=ResNet-502023.11 | 55.7 | |
| TimeSformer KineticsPrivacy Preserving=false, Stage 2: Alignment=Kinetics, Evaluation Protocol=Linear Probing, Backbone=ViT-B, Pre-trained on ImageNet-21K=true2023.11 | 55.4 | |
| CoCLRModality=V2023.03 | 54.6 | |
| MemDPCModality=V2023.03 | 54.5 | |
| TimeSformer SyntheticPrivacy Preserving=false, Stage 2: Alignment=Synthetic, Evaluation Protocol=Fine-tuning, Backbone=ViT-B, Pre-trained on ImageNet-21K=true2023.11 | 54.4 | |
| TSN (RN50 backbone)Privacy Preserving=true, Stage 2: Alignment=Synthetic, Evaluation Protocol=Fine-tuning, Backbone=ResNet-502023.11 | 54.4 | |
| InternVideo2_clip-1B#F=8, Training Data=IV-25.5M, protocol=Zero-shot2024.03 | 53.9 | |
| R(2+1)D (RN50 backbone)Privacy Preserving=true, Stage 2: Alignment=Synthetic, Evaluation Protocol=Fine-tuning, Backbone=ResNet-502023.11 | 53.3 | |
| TVTSv2#F=12, Training Data=V-8.5M, protocol=Zero-shot2024.03 | 52.1 | |
| Hierarchical Clustering Multi-Task LearningYear=20172025.01 | 51.41 | |
| VicTRYear=2024, Backbone=ViT-B/162025.01 | 51 | |
| VTHCLModality=V2023.03 | 49.2 | |
| VideoMoCoModality=V2023.03 | 49.2 | |
| TimeSformer SyntheticPrivacy Preserving=false, Stage 2: Alignment=Synthetic, Evaluation Protocol=Linear Probing, Backbone=ViT-B, Pre-trained on ImageNet-21K=true2023.11 | 49.2 | |
| SpeedNetModality=V2023.03 | 48.8 | |
| CLIP#F=12, Training Data=I-400M, protocol=Zero-shot2024.03 | 43.2 | |
| PaceModality=V2023.03 | 36.6 | |
| VCOPModality=V2023.03 | 30.9 | |
| OPNModality=V2023.03 | 23.8 | |
| I3D (RN50 backbone)Privacy Preserving=true, Stage 2: Alignment=Synthetic, Evaluation Protocol=Linear Probing, Backbone=ResNet-502023.11 | 22.6 | |
| R(2+1)D (RN50 backbone)Privacy Preserving=true, Stage 2: Alignment=Synthetic, Evaluation Protocol=Linear Probing, Backbone=ResNet-502023.11 | 22.2 | |
| TSN (RN50 backbone)Privacy Preserving=true, Stage 2: Alignment=Synthetic, Evaluation Protocol=Linear Probing, Backbone=ResNet-502023.11 | 20.9 |