Action Recognition on UAV-Human
77.4Top-1 AccFALCON
Evaluation Results
| Method | Links | |
|---|---|---|
| FALCONBackbone=ViT-L, Extra data=K400, Input Size=224 × 224, Frames=16, GFLOPs=597, Params=305M2024.09 | 77.4 | |
| VideoMAEBackbone=ViT-L, Extra data=K400, Input Size=224 × 224, Frames=16, GFLOPs=597, Params=305M2024.09 | 71.5 | |
| FALCONBackbone=ViT-B, Extra data=K400, Input Size=224 × 224, Frames=16, GFLOPs=180, Params=87M2024.09 | 67.9 | |
| VideoMAEBackbone=ViT-B, Extra data=K400, Input Size=224 × 224, Frames=16, GFLOPs=180, Params=87M2024.09 | 62.1 | |
| VideoMAE v2Backbone=ViT-G, Extra data=K400, Input Size=224 × 224, Frames=16, GFLOPs=5100, Params=632M2024.09 | 61.1 | |
| MVDBackbone=ViT-B, Extra data=IN21K + K400, Input Size=224 × 224, Frames=16, GFLOPs=180, Params=87M2024.09 | 60.5 | |
| SiamMAEBackbone=ViT-B, Extra data=K400, Input Size=224 × 224, Frames=16, GFLOPs=180, Params=87M2024.09 | 60.1 | |
| PMI SamplerBackbone=X3D-M, Extra data=K400, Input Size=224 × 224, Frames=16, GFLOPs=7, Params=4M2024.09 | 55 | |
| MITFASBackbone=X3D-M, Extra data=K400, Input Size=224 × 224, Frames=16, GFLOPs=7, Params=4M2024.09 | 50.8 | |
| MotionFormerBackbone=ViT-B, Extra data=IN21K + K400, Input Size=224 × 224, Frames=8, GFLOPs=370, Params=109M2024.09 | 50.4 | |
| AZTRBackbone=X3D-M, Extra data=K400, Input Size=224 × 224, Frames=16, GFLOPs=7, Params=4M2024.09 | 47.4 | |
| ST-MAEBackbone=ViT-B, Extra data=K400, Input Size=224 × 224, Frames=16, GFLOPs=180, Params=87M2024.09 | 45.1 | |
| DiffFARBackbone=X3D-M, Extra data=K400, Input Size=540 × 960, Frames=8, GFLOPs=130, Params=4M2024.09 | 41.9 | |
| FARBackbone=X3D-M, Extra data=K400, Input Size=540 × 960, Frames=8, GFLOPs=65, Params=4M2024.09 | 38.6 | |
| Privacy-Preserving MAE-Align (PPMA)Privacy Preserving=true, Stage 1: MAE=NH Kinetics, Stage 2: Alignment=NH Kinetics + Synthetic, Evaluation Protocol=Fine-tuning, Backbone=ViT-B2023.11 | 38.5 | |
| TimesFormerBackbone=ViT-B, Extra data=K400, Input Size=224 × 224, Frames=8, GFLOPs=196, Params=131M2024.09 | 38.4 | |
| MAE-Align w/ SyntheticPrivacy Preserving=true, Stage 1: MAE=Synthetic, Stage 2: Alignment=Synthetic, Evaluation Protocol=Fine-tuning, Backbone=ViT-B2023.11 | 36.1 | |
| TSN (RN50 backbone)Privacy Preserving=true, Stage 2: Alignment=Synthetic, Evaluation Protocol=Fine-tuning, Backbone=ResNet-502023.11 | 35.6 | |
| I3D (RN50 backbone)Privacy Preserving=true, Stage 2: Alignment=Synthetic, Evaluation Protocol=Fine-tuning, Backbone=ResNet-502023.11 | 35.1 | |
| MAE-Align w/ real humansPrivacy Preserving=false, Stage 1: MAE=Kinetics, Stage 2: Alignment=Kinetics, Evaluation Protocol=Fine-tuning, Backbone=ViT-B2023.11 | 34.9 | |
| ViViT FEBackbone=ViT-B, Extra data=IN-21K, Input Size=224 × 224, Frames=16, GFLOPs=284, Params=116M2024.09 | 34.1 | |
| R(2+1)D (RN50 backbone)Privacy Preserving=true, Stage 2: Alignment=Synthetic, Evaluation Protocol=Fine-tuning, Backbone=ResNet-502023.11 | 31.8 | |
| TimeSformer SyntheticPrivacy Preserving=false, Stage 2: Alignment=Synthetic, Evaluation Protocol=Fine-tuning, Backbone=ViT-B, Pre-trained on ImageNet-21K=true2023.11 | 25 | |
| MViT v1Backbone=MViT-B, Extra data=K400, Input Size=224 × 224, Frames=16, GFLOPs=71, Params=37M2024.09 | 24.3 | |
| TimeSformer KineticsPrivacy Preserving=false, Stage 2: Alignment=Kinetics, Evaluation Protocol=Fine-tuning, Backbone=ViT-B, Pre-trained on ImageNet-21K=true2023.11 | 23.3 | |
| MAE-Align w/ SyntheticPrivacy Preserving=true, Stage 1: MAE=Synthetic, Stage 2: Alignment=Synthetic, Evaluation Protocol=Linear Probing, Backbone=ViT-B2023.11 | 20.6 | |
| Privacy-Preserving MAE-Align (PPMA)Privacy Preserving=true, Stage 1: MAE=NH Kinetics, Stage 2: Alignment=NH Kinetics + Synthetic, Evaluation Protocol=Linear Probing, Backbone=ViT-B2023.11 | 19.3 | |
| MAE-Align w/ real humansPrivacy Preserving=false, Stage 1: MAE=Kinetics, Stage 2: Alignment=Kinetics, Evaluation Protocol=Linear Probing, Backbone=ViT-B2023.11 | 13.9 | |
| TimeSformer SyntheticPrivacy Preserving=false, Stage 2: Alignment=Synthetic, Evaluation Protocol=Linear Probing, Backbone=ViT-B, Pre-trained on ImageNet-21K=true2023.11 | 13.8 | |
| I3D (RN50 backbone)Privacy Preserving=true, Stage 2: Alignment=Synthetic, Evaluation Protocol=Linear Probing, Backbone=ResNet-502023.11 | 5.8 | |
| TSN (RN50 backbone)Privacy Preserving=true, Stage 2: Alignment=Synthetic, Evaluation Protocol=Linear Probing, Backbone=ResNet-502023.11 | 5.7 | |
| R(2+1)D (RN50 backbone)Privacy Preserving=true, Stage 2: Alignment=Synthetic, Evaluation Protocol=Linear Probing, Backbone=ResNet-502023.11 | 5.5 | |
| TimeSformer KineticsPrivacy Preserving=false, Stage 2: Alignment=Kinetics, Evaluation Protocol=Linear Probing, Backbone=ViT-B, Pre-trained on ImageNet-21K=true2023.11 | 5.3 | |
| ScratchPrivacy Preserving=false, Evaluation Protocol=Fine-tuning, Backbone=ViT-B2023.11 | 0.7 |