Action Recognition on NTU RGB+D 60 (X-View)
99.6AccuracyPoseC3D
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| PoseC3DModality=RGB + Estimated Pose, Pre-trained=true2022.12 | 99.6 | — | |
| PoseConv3DModality=B+RGB, GFLOPs=41.8*, Params(M)=31.6*2026.01 | 99.6 | — | |
| SkeleTGFLOPs=9.6, Params(M)=5.22026.01 | 99.6 | — | |
| BHaRNet-MModality=J+B+JM+BM+RGB, GFLOPs=13.0, Params(M)=22.22026.01 | 99.4 | — | |
| BHaRNet-MModality=J+B+RGB, GFLOPs=7.6, Params(M)=16.72026.01 | 99.3 | — | |
| Ours (BHaRNet-E)GFLOPs=10.9, Params(M)=112026.01 | 99.2 | — | |
| Ours (BHaRNet-P)GFLOPs=6.8, Params(M)=9.72026.01 | 99.2 | — | |
| VPN++ (3D pose)Modality=RGB + Pose, Pre-trained=true2022.12 | 99.1 | — | |
| MMNetModality=J+B+RGB, GFLOPs=89.2, Params(M)=34.22026.01 | 99.1 | — | |
| BHaRNet-P‡GFLOPs=3.4, Params(M)=4.92026.01 | 99.1 | — | |
| Ours (BHaRNet-B)GFLOPs=6.6, Params(M)=5.52026.01 | 99.1 | — | |
| EPAM-NetModality=J+RGB, GFLOPs=8.1, Params(M)=2.52026.01 | 99 | — | |
| BHaRNet-M†Modality=J+B+RGB, GFLOPs=7.6, Params(M)=16.72026.01 | 99 | — | |
| BHaRNet-B‡GFLOPs=9.6, Params(M)=5.22026.01 | 99 | — | |
| BHaRNet-E‡GFLOPs=3.3, Params(M)=2.82026.01 | 99 | — | |
| MMNetModality=RGB + Pose, Pre-trained=true2022.12 | 98.8 | — | |
| BHaRNet-E†GFLOPs=5.4, Params(M)=5.52026.01 | 98.8 | — | |
| BHaRNet-P†GFLOPs=3.4, Params(M)=4.92026.01 | 98.8 | — | |
| 3MformerGFLOPs=58.5, Params(M)=6.72026.01 | 98.7 | — | |
| BHaRNet-B†GFLOPs=3.3, Params(M)=2.82026.01 | 98.7 | — | |
| LLM-ARCategory=LLM, Mod.=J, Two-stage=Y, Interpret.=Y2026.05 | 98.4 | — | |
| REASON+LLMCategory=LLM, Mod.=J, Two-stage=N, Interpret.=Y2026.05 | 98.3 | — | |
| VPNModality=RGB + Pose, Pre-trained=true2022.12 | 98 | — | |
| OursModality=RGB + Pose, Pre-trained=false2022.12 | 97.9 | — | |
| pi-ViTModality=J+RGB, GFLOPs=590.0, Params(M)=121.42026.01 | 97.9 | — | |
| ProtoGCNEnsemble size=62024.11 | 97.8 | — | |
| ProtoGCNGFLOPs=43.4, Params(M)=24.92026.01 | 97.8 | — | |
| Hyper-GCNCategory=GCN, Mod.=J+B+JM+BM, Two-stage=N, Interpret.=N2026.05 | 97.8 | — | |
| REASONCategory=GCN, Mod.=J+B+JM+BM, Two-stage=N, Interpret.=Y2026.05 | 97.8 | — | |
| SkateFormerCategory=Trans., Mod.=J+B+JM+BM, Two-stage=N, Interpret.=N2026.05 | 97.8 | — | |
| SUGARCategory=LLM, Mod.=J, Two-stage=Y, Interpret.=Y2026.05 | 97.8 | — | |
| MAMPBackbone=Transformer, evaluation protocol=fine-tuned2023.08 | 97.5 | — | |
| JT-GraphFormer2024.11 | 97.5 | — | |
| DS-GCN2024.11 | 97.5 | — | |
| ProtoGCNEnsemble size=42024.11 | 97.5 | — | |
| DS-GCNCategory=GCN, Mod.=J+B+JM+BM, Two-stage=N, Interpret.=N2026.05 | 97.5 | — | |
| TSMFModality=RGB + Pose, Pre-trained=false2022.12 | 97.4 | — | |
| TSMFTraining Modality=Multi-modality, Inference Modality=Multi-modality, Multi-stream ensemble=true2024.07 | 97.4 | — | |
| MMCLTraining Modality=Multi-modality, Inference Modality=Pose, Multi-stream ensemble=true2024.07 | 97.4 | — | |
| PYSKL2024.11 | 97.4 | — | |
| DeGCNGFLOPs=6.9, Params(M)=5.62026.01 | 97.4 | — | |
| HD-GCN2024.11 | 97.2 | — | |
| ProtoGCNEnsemble size=22024.11 | 97.2 | — | |
| InfoGCNModality=Pose, Pre-trained=false2022.12 | 97.1 | — | |
| ACFL-CTRTraining Modality=Pose, Inference Modality=Pose, Multi-stream ensemble=true2024.07 | 97.1 | — | |
| InfoGCNTraining Modality=Pose, Inference Modality=Pose, Multi-stream ensemble=true2024.07 | 97.1 | — | |
| InfoGCN2024.11 | 97.1 | — | |
| SkeletonGCL2024.11 | 97.1 | — | |
| InfoGCNGFLOPs=10.0*, Params(M)=9.42026.01 | 97.1 | — | |
| PoseConv3DGFLOPs=31.8, Params(M)=42026.01 | 97.1 | — | |
| SkeletonGCLTraining Modality=Pose, Inference Modality=Pose, Multi-stream ensemble=true2024.07 | 97 | — | |
| LSTTraining Modality=Multi-modality, Inference Modality=Pose, Multi-stream ensemble=true2024.07 | 97 | — | |
| GAP2024.11 | 97 | — | |
| BlockGCN2024.11 | 97 | — | |
| BlockGCNGFLOPs=6.5, Params(M)=5.22026.01 | 97 | — | |
| HD-GCN*Category=GCN, Mod.=J+B+J'+B', Two-stage=N, Interpret.=N2026.05 | 97 | — | |
| BlockGCNCategory=GCN, Mod.=J+B+JM+BM, Two-stage=N, Interpret.=N2026.05 | 97 | — | |
| GAPCategory=LLM, Mod.=J+B+JM+BM, Two-stage=N, Interpret.=N2026.05 | 97 | — | |
| STF2024.11 | 96.9 | — | |
| InfoGCNCategory=GCN, Mod.=J+B+JM+BM, Two-stage=N, Interpret.=N2026.05 | 96.9 | — | |
| CTR-GCNModality=Pose, Pre-trained=false2022.12 | 96.8 | — | |
| CTR-GCNTraining Modality=Pose, Inference Modality=Pose, Multi-stream ensemble=true2024.07 | 96.8 | — | |
| SAP-CTRTraining Modality=Pose, Inference Modality=Pose, Multi-stream ensemble=true2024.07 | 96.8 | — | |
| FR-HeadTraining Modality=Pose, Inference Modality=Pose, Multi-stream ensemble=true2024.07 | 96.8 | — | |
| KoopmanTraining Modality=Pose, Inference Modality=Pose, Multi-stream ensemble=true2024.07 | 96.8 | — | |
| CTR-GCN2024.11 | 96.8 | — | |
| FR-Head2024.11 | 96.8 | — | |
| CTR-GCNGFLOPs=7.9, Params(M)=5.82026.01 | 96.8 | — | |
| FR HeadCategory=GCN, Mod.=J+B+JM+BM, Two-stage=N, Interpret.=N2026.05 | 96.8 | — | |
| DST-HCNCategory=GCN, Mod.=J+B+JM+BM, Two-stage=N, Interpret.=N2026.05 | 96.8 | — | |
| Tripoolinference_mode=clip-based, streams=32022.03 | 96.7 | — | |
| STH-DRLinference_mode=clip-based, streams=12022.03 | 96.7 | — | |
| Skeletal GNNModality=Pose, Pre-trained=false2022.12 | 96.7 | — | |
| PSUMNetTraining Modality=Pose, Inference Modality=Pose, Multi-stream ensemble=true2024.07 | 96.7 | — | |
| PoseC3DModality=Estimated Pose, Pre-trained=true2022.12 | 96.6 | — | |
| DualHead-NetModality=Pose, Pre-trained=false2022.12 | 96.6 | — | |
| MST-GCNTraining Modality=Pose, Inference Modality=Pose, Multi-stream ensemble=true2024.07 | 96.6 | — | |
| MG-GCNTraining Modality=Pose, Inference Modality=Pose, Multi-stream ensemble=true2024.07 | 96.6 | — | |
| DC-GCN+ADG2024.11 | 96.6 | — | |
| MST-GCN2024.11 | 96.6 | — | |
| DC-GCN+ADGCategory=GCN, Mod.=J+B+JM+BM, Two-stage=N, Interpret.=N2026.05 | 96.6 | — | |
| MST-GCNCategory=GCN, Mod.=J+B+JM+BM, Two-stage=N, Interpret.=N2026.05 | 96.6 | — | |
| Selective-HCNCategory=GCN, Mod.=J+B+JM+BM, Two-stage=N, Interpret.=N2026.05 | 96.6 | — | |
| STF-Netinference_mode=clip-based, streams=42022.03 | 96.5 | — | |
| ShiftGCNinference_mode=clip-based, streams=42022.03 | 96.5 | 10 | |
| STARModality=RGB + Pose, Pre-trained=false2022.12 | 96.5 | — | |
| Shift-GCNTraining Modality=Pose, Inference Modality=Pose, Multi-stream ensemble=true2024.07 | 96.5 | — | |
| Tripoolinference_mode=clip-based, streams=22022.03 | 96.4 | — | |
| ViABackbone=2s-UINK, evaluation protocol=fine-tuned2023.08 | 96.4 | — | |
| CTR-GCNCategory=GCN, Mod.=J+B+JM+BM, Two-stage=N, Interpret.=N2026.05 | 96.4 | — | |
| ST-TRinference_mode=clip-based, streams=22022.03 | 96.3 | — | |
| ShiftGCN++inference_mode=clip-based, streams=42022.03 | 96.3 | 1.7 | |
| MCCBackbone=2s-AGCN, evaluation protocol=fine-tuned2023.08 | 96.3 | — | |
| MS-G3Dinference_mode=clip-based, streams=22022.03 | 96.2 | — | |
| MS-AAGCNinference_mode=clip-based, streams=42022.03 | 96.2 | — | |
| STF-Netinference_mode=clip-based, streams=22022.03 | 96.2 | — | |
| MS-G3DTraining Modality=Pose, Inference Modality=Pose, Multi-stream ensemble=true2024.07 | 96.2 | — | |
| VPNTraining Modality=Multi-modality, Inference Modality=Multi-modality, Multi-stream ensemble=true2024.07 | 96.2 | — | |
| MS-G3D2024.11 | 96.2 | — | |
| MS-G3DCategory=GCN, Mod.=J+B+JM+BM, Two-stage=N, Interpret.=N2026.05 | 96.2 | — |