Action Recognition on Toyota SmartHome (TSH) (CV2)
73.3AccuracyPAN-Ensemble
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| PAN-EnsembleModality - 2D Ske=false, Modality - 3D Ske=true, Modality - RGB=true, Modeling Paradigm=Multimodal (separate modeling)2025.12 | 73.3 | — | — | |
| PAN (Even)Modality - 2D Ske=false, Modality - 3D Ske=false, Modality - RGB=true, Modeling Paradigm=Unimodal, Sampling Strategy=token replication2025.12 | 72.9 | — | — | |
| PAN-UnifiedModality - 2D Ske=false, Modality - 3D Ske=true, Modality - RGB=true, Modeling Paradigm=Multimodal (unified modeling)2025.12 | 72.1 | — | — | |
| PAN (Guided)Modality - 2D Ske=true, Modality - 3D Ske=false, Modality - RGB=true, Modeling Paradigm=Multimodal (separate modeling)2025.12 | 71.1 | — | — | |
| ViAPose=Yes, RGB=No2026.03 | 65.4 | — | — | |
| Proposed Multi-modal ArchitecturePose=Yes, RGB=Yes2026.03 | 65.4 | — | — | |
| UNIKRGB=false, Pose=true, Pre-training=Posetics(Ours)2021.07 | 65 | — | — | |
| π-ViT + 3D PosesPose=Yes, RGB=Yes2023.11 | 65 | — | — | |
| π-ViTModality - 2D Ske=training only, Modality - 3D Ske=true, Modality - RGB=true, Modeling Paradigm=Multimodal (separate modeling), Fusion=late fusion2025.12 | 65 | — | — | |
| π-ViT + 3D PosesPose=Yes, RGB=Yes2026.03 | 65 | — | — | |
| π-ViTPose=Only during training, RGB=Yes2023.11 | 64.8 | — | — | |
| π-ViTModality - 2D Ske=training only, Modality - 3D Ske=training only, Modality - RGB=true, Modeling Paradigm=Multimodal (separate modeling)2025.12 | 64.8 | — | — | |
| π-ViTPose=Training Only, RGB=Yes2026.03 | 64.8 | — | — | |
| TimeSformer + 2D-SIMPose=Only during training, RGB=Yes2023.11 | 62.9 | — | — | |
| TimeSformer + 3D-SIMPose=Only during training, RGB=Yes2023.11 | 62.3 | — | — | |
| UNIKPose=Yes, RGB=No2026.03 | 61.4 | — | — | |
| UNIKRGB=false, Pose=true, Pre-training=Scratch2021.07 | 61.2 | — | — | |
| TimeSformerPose=No, RGB=Yes2023.11 | 60.6 | — | — | |
| TimeSformerModality - 2D Ske=false, Modality - 3D Ske=false, Modality - RGB=true, Modeling Paradigm=Unimodal2025.12 | 60.6 | — | — | |
| MS-G3D NetRGB=false, Pose=true, Pre-training=Scratch2021.07 | 59.4 | — | — | |
| MS-G3D NetPose=Yes, RGB=No2026.03 | 59.4 | — | — | |
| PAN (Even)Modality - 2D Ske=false, Modality - 3D Ske=false, Modality - RGB=true, Modeling Paradigm=Unimodal, Sampling Strategy=zero padding2025.12 | 58.7 | — | — | |
| VPN++ + 3D PosesPose=true, RGB=true, Att=true, Pose Quality=New Poses2021.05 | 58.1 | — | — | |
| VPN++ + 3D PosesPose=Yes, RGB=Yes2023.11 | 58.1 | — | — | |
| VPN++Modality - 2D Ske=false, Modality - 3D Ske=true, Modality - RGB=true, Modeling Paradigm=Multimodal (separate modeling), Fusion=late fusion2025.12 | 58.1 | — | — | |
| VPN++ + 3D PosesPose=Yes, RGB=Yes2026.03 | 58.1 | — | — | |
| SV-data2vecPose=Training Only, RGB=Yes2026.03 | 57.5 | — | — | |
| VPN++Pose=training only, RGB=true, Att=true, Pose Quality=New Poses2021.05 | 54.9 | — | — | |
| VPN++Pose=Only during training, RGB=Yes2023.11 | 54.9 | — | — | |
| VPN++Modality - 2D Ske=false, Modality - 3D Ske=training only, Modality - RGB=true, Modeling Paradigm=Multimodal (separate modeling)2025.12 | 54.9 | — | — | |
| VPN++Pose=Training Only, RGB=Yes2026.03 | 54.9 | — | — | |
| Neural Projection Layer2021.03 | 54.6 | — | — | |
| NPLRGB=true, Pose=false, Pre-training=Kinetics-4002021.07 | 54.6 | — | — | |
| LTNPose=No, RGB=Yes2023.11 | 54.6 | — | — | |
| LTNPose=No, RGB=Yes2026.03 | 54.6 | — | — | |
| VPNPose=true, RGB=true, Att=true, Pose Quality=New Poses2021.05 | 54.1 | — | — | |
| VPNPose=Yes, RGB=Yes2023.11 | 54.1 | — | — | |
| VPNPose=Yes, RGB=Yes2026.03 | 54.1 | — | — | |
| VPN++Pose=training only, RGB=true, Att=true, Pose Quality=Old Poses2021.05 | 53.6 | — | — | |
| VPNPose=true, RGB=true, Att=true, Pose Quality=Old Poses2021.05 | 53.5 | — | — | |
| VPNRGB=true, Pose=false, Pre-training=Kinetics-4002021.07 | 53.5 | — | — | |
| 2s-AGCNRGB=false, Pose=true, Pre-training=Scratch2021.07 | 53.5 | — | — | |
| ST-GCNRGB=false, Pose=true, Pre-training=Scratch2021.07 | 51.1 | — | — | |
| MotionFormerPose=No, RGB=Yes2023.11 | 51 | — | — | |
| MotionFormerModality - 2D Ske=false, Modality - 3D Ske=false, Modality - RGB=true, Modeling Paradigm=Unimodal2025.12 | 51 | — | — | |
| ST-GCNPose=Yes, RGB=No2026.03 | 51 | — | — | |
| STA2021.03 | 50.3 | — | — | |
| P-I3DPose=true, RGB=true, Att=true, Pose Quality=Old Poses2021.05 | 50.3 | — | — | |
| Separable STAPose=true, RGB=true, Att=true, Pose Quality=Old Poses2021.05 | 50.3 | — | — | |
| Separable STARGB=true, Pose=false, Pre-training=Kinetics-4002021.07 | 50.3 | — | — | |
| P-I3DPose=Yes, RGB=Yes2023.11 | 50.3 | — | — | |
| Separable STAPose=Yes, RGB=Yes2023.11 | 50.3 | — | — | |
| P-I3DPose=Yes, RGB=Yes2026.03 | 50.3 | — | — | |
| Separable STAPose=Yes, RGB=Yes2026.03 | 50.3 | — | — | |
| 2s-AGCNPose=true, RGB=false, Att=false, Pose Quality=New Poses2021.05 | 49.7 | — | — | |
| Video SwinPose=No, RGB=Yes2023.11 | 48.6 | — | — | |
| Video SwinModality - 2D Ske=false, Modality - 3D Ske=false, Modality - RGB=true, Modeling Paradigm=Unimodal2025.12 | 48.6 | — | — | |
| MMNetPose=Yes, RGB=Yes2026.03 | 46.6 | — | — | |
| I3D2021.03 | 45.1 | — | — | |
| I3DPose=false, RGB=true, Att=false2021.05 | 45.1 | — | — | |
| I3DRGB=true, Pose=false, Pre-training=Kinetics-4002021.07 | 45.1 | — | — | |
| I3DPose=No, RGB=Yes2026.03 | 45.1 | — | — | |
| I3D+NLPose=false, RGB=true, Att=true2021.05 | 43.9 | — | — | |
| HyperformerPose=Yes, RGB=No2023.11 | 35.2 | — | — | |
| HyperformerModality - 2D Ske=false, Modality - 3D Ske=true, Modality - RGB=false, Modeling Paradigm=Unimodal2025.12 | 35.2 | — | — | |
| PoseC3DPose=Yes, RGB=Yes2023.11 | 33.4 | — | — | |
| PoseC3DModality - 2D Ske=re-estimated data, Modality - 3D Ske=false, Modality - RGB=true, Modeling Paradigm=Multimodal (separate modeling)2025.12 | 33.4 | — | — | |
| 2s-AGCNPose=Yes, RGB=No2023.11 | 32.3 | — | — | |
| 2s-AGCNModality - 2D Ske=false, Modality - 3D Ske=true, Modality - RGB=false, Modeling Paradigm=Unimodal2025.12 | 32.3 | — | — | |
| 2s-AGCNPose=Yes, RGB=No2026.03 | 32.3 | — | — | |
| PoseC3DPose=Yes, RGB=No2023.11 | 28.2 | — | — | |
| PoseC3DModality - 2D Ske=re-estimated data, Modality - 3D Ske=false, Modality - RGB=false, Modeling Paradigm=Unimodal2025.12 | 28.2 | — | — | |
| IDT2021.03 | 23.7 | — | — | |
| DTPose=false, RGB=true, Att=false2021.05 | 23.7 | — | — | |
| Pose LSTM2021.03 | 17.2 | — | — | |
| LSTMPose=true, RGB=false, Att=false, Pose Quality=Old Poses2021.05 | 17.2 | — | — | |
| LSTMRGB=false, Pose=true, Pre-training=Scratch2021.07 | 17.2 | — | — | |
| EPAM-NetSkeleton Modality=true, RGB Modality=true2024.08 | — | 67.8 | — | |
| MMNet (ResNet18)Skeleton Modality=true, RGB Modality=true, backbone=ResNet182024.08 | — | 33.4 | — | |
| TimeSformerBackbone=TimeSformer2024.06 | — | 36.6 | — | |
| TimeSformerPre-train=Kinetics400, Resolution=224x224, Frames=82026.06 | — | — | 59.5 | |
| TimeSformer + BigBirdBackbone=TimeSformer, Attention Mechanism=BigBird2024.06 | — | 40.1 | — | |
| TimeSformer + Diff. Attn.Pre-train=Kinetics400, Resolution=224x224, Frames=82026.06 | — | — | 59.4 | |
| TimeSformer + DnAPre-train=Kinetics400, Resolution=224x224, Frames=82026.06 | — | — | 63.5 | |
| TimeSformer + Fibottention (Modified Wythoff)Backbone=TimeSformer, Instantiation=Modified Wythoff2024.06 | — | 42.3 | — | |
| TimeSformer + Fibottention (Wythoff)Backbone=TimeSformer, Instantiation=Wythoff2024.06 | — | 38.6 | — | |
| TSMFSkeleton Modality=true, RGB Modality=true2024.08 | — | 28.9 | — | |
| VPN (I3D)Skeleton Modality=true, RGB Modality=true, backbone=I3D2024.08 | — | 54.1 | — | |
| VPN++Skeleton Modality=false, RGB Modality=true, Skeleton used in training but not in inference=true2024.08 | — | 54.9 | — | |
| VPN+++ 3D PosesSkeleton Modality=true, RGB Modality=true2024.08 | — | 58.1 | — | |
| π-ViTSkeleton Modality=false, RGB Modality=true, 10 clips per video=true, Skeleton used in training but not in inference=true2024.08 | — | 64.8 | — | |
| π-ViT + 3D Poses+Skeleton Modality=true, RGB Modality=true, 10 clips per video=true2024.08 | — | 65 | — |