Action Recognition on NTU RGB+D 120 (XSub & XSet)
92.9Top-1 Accuracy (XSub)MMNet (Inception-v3)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| MMNet (Inception-v3)Modality=Multimodal (Skeleton+RGB), Backbone=Inception-v3, GFLOPs=89.2, Param (M)=34.22024.08 | 92.9 | 94.4 | |
| EPAM-Net (Our proposed approach)Modality=Multimodal (Skeleton+RGB), GFLOPs=8.1, Param (M)=2.52024.08 | 92.4 | 94.3 | |
| π-ViTModality=RGB, Skeleton usage=Training only, Input clips per video=10, GFLOPs=590.0, Param (M)=121.42024.08 | 91.9 | 92.9 | |
| VPN++ + 3D PosesModality=Multimodal (Skeleton+RGB), GFLOPs=125.8, Param (M)=15.52024.08 | 90.7 | 92.5 | |
| TSMFModality=Multimodal (Skeleton+RGB), GFLOPs=85.4, Param (M)=20.82024.08 | 87 | 89.1 | |
| MS-G3DModality=Skeleton, GFLOPs=16.7, Param (M)=2.82024.08 | 86.9 | 88.4 | |
| VPN++Modality=RGB, Skeleton usage=Training only2024.08 | 86.7 | 89.3 | |
| VPN (I3D)Modality=Multimodal (Skeleton+RGB), Backbone=I3D, GFLOPs=107.9, Param (M)=24.02024.08 | 86.3 | 87.8 | |
| PoseConv3DModality=Skeleton, GFLOPs=15.9, Param (M)=2.02024.08 | 86 | 89.6 | |
| 2s-AGCNModality=Skeleton, GFLOPs=8.8, Param (M)=3.52024.08 | 82.9 | 84.9 | |
| ST-GCNModality=Skeleton, GFLOPs=3.8, Param (M)=3.12024.08 | 79 | 81.3 |