Action Recognition on Diving-48
94.9Top-1 AccLVMAE (ViT-L)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| LVMAE (ViT-L)Pre-train=Unlabeled K7102024.11 | 94.9 | — | |
| LVMAE (ViT-B)Pre-train=Unlabeled K7102024.11 | 91.2 | — | |
| MC-ViT-LPre-train=ALIGN+LTIP+JFT+HT100M+VTP2024.11 | 91 | — | |
| Video-FocalNet-BPre-train=K4002024.11 | 90.8 | — | |
| AIM ViT-L/14Pre-train=CLIP2024.11 | 90.6 | — | |
| V-JEPA 2 ViT-gParam.=1B, Resolution=384x384, Frame clips=32, Temporal crops=4, Spatial crops=32026.03 | 90.2 | — | |
| VJEPA2Type=Discriminative Pretraining, Pretrained on=Videos, Arch.=ViT-G, Param.=1B, Attentive probing=true2025.12 | 90.1 | — | |
| MC-ViT-BPre-train=ALIGN+LTIP+JFT+HT100M+VTP2024.11 | 89.7 | — | |
| V-JEPA 2.1 ViT-GParam.=2B, Resolution=384x384, Frame clips=32, Temporal crops=4, Spatial crops=32026.03 | 89.2 | — | |
| V-JEPA 2.1 ViT-gParam.=1B, Resolution=384x384, Frame clips=32, Temporal crops=4, Spatial crops=32026.03 | 89 | — | |
| AIM ViT-B/16Pre-train=CLIP2024.11 | 88.9 | — | |
| ORVIT TimeSformerPretrain=IN-21K, Frames=32, Params (10^6)=160 (+32%)2021.10 | 88 | — | |
| ORVITPre-train=IN21K2024.11 | 88 | — | |
| VJEPAType=Discriminative Pretraining, Pretrained on=Videos, Arch.=ViT-H, Param.=630M, Attentive probing=true2025.12 | 87.9 | — | |
| V-JEPA ViT-HParam.=600M, Resolution=256x256, Frame clips=32, Temporal crops=4, Spatial crops=32026.03 | 87.9 | — | |
| LVMAE (ViT-B)Pre-train=None2024.11 | 87.8 | — | |
| SIFAR-B-14Pre-train=IN21K2024.11 | 87.3 | — | |
| BEVTPretrain=IN-1K+K400, Params=88.1M, Tokenizer=PeCo2021.12 | 87.2 | — | |
| NExT-Vid-GType=Generative Pretraining, Pretrained on=Videos, Arch.=ViT-G, Param.=1.1B, Attentive probing=true2025.12 | 87.2 | — | |
| ORVIT TimeSformer-LightPretrain=IN-21K, Frames=32, Params (10^6)=126 (+3%)2021.10 | 86.8 | — | |
| BEVTPretrain=IN-1K+K400, Params=88.1M, Tokenizer=DALL-E2021.12 | 86.7 | — | |
| BEVTPre-train=IN21K+K4002024.11 | 86.7 | — | |
| InternVideo2Type=Discriminative Pretraining, Pretrained on=Videos, Arch.=ViT-G, Param.=1B, Attentive probing=true2025.12 | 86.4 | — | |
| InternVideo2s2-1BParam.=1B, Resolution=256x256, Frame clips=32, Temporal crops=4, Spatial crops=32026.03 | 86.4 | — | |
| NExT-Vid-HType=Generative Pretraining, Pretrained on=Videos, Arch.=ViT-H, Param.=600M, Attentive probing=true2025.12 | 84.5 | — | |
| RSANet-R50FLOPs x clips=72 G x 22021.11 | 84.2 | — | |
| Swin-BPretrain=IN-1K, Params=88.1M2021.12 | 84 | — | |
| TimeSformer† + STRG + STINPretrain=IN-21K, Frames=32, Params (10^6)=1322021.10 | 83.5 | — | |
| NExT-Vid-LType=Generative Pretraining, Pretrained on=Videos, Arch.=ViT-L, Param.=300M, Attentive probing=true2025.12 | 82.7 | — | |
| DINOv2Type=Discriminative Pretraining, Pretrained on=Images, Arch.=ViT-G, Param.=1.1B, Attentive probing=true2025.12 | 82.5 | — | |
| DINOv2Param.=1.1B, Resolution=256x256, Frame clips=32, Temporal crops=4, Spatial crops=32026.03 | 82.5 | — | |
| VideoSwin-BPre-train=IN21K2024.11 | 81.9 | — | |
| TQNPretrain=K400, Params=N/A2021.12 | 81.8 | — | |
| TQNPretrain=K400, Frames=ALL2021.10 | 81.8 | — | |
| PMI SamplerFrames=16, Backbone=X3D-M2023.04 | 81.3 | — | |
| TimeSformer-LFLOPs x clips=2380 G x 32021.11 | 81 | — | |
| TimeSformer-LPretrain=IN-21K, Params=121.4M2021.12 | 81 | — | |
| TimeSformer-L†Pretrain=IN-21K, Frames=96, Params (10^6)=1212021.10 | 81 | — | |
| TimeSformer† + STINPretrain=IN-21K, Frames=32, Params (10^6)=1232021.10 | 81 | — | |
| TimeSformer-LPre-train=IN21K2024.11 | 81 | — | |
| TimeSformer†Pretrain=IN-21K, Frames=32, Params (10^6)=1212021.10 | 80 | — | |
| GST-50†Pretrain=ImageNet, Frames=82021.10 | 78.9 | — | |
| TimeSformer† + STRGPretrain=IN-21K, Frames=32, Params (10^6)=1292021.10 | 78.1 | — | |
| TimeSformer-HRFLOPs x clips=1703 G x 32021.11 | 78 | — | |
| TimeSformer-HR†Pretrain=IN-21K, Frames=16, Params (10^6)=1212021.10 | 78 | — | |
| SlowFast-R101FLOPs x clips=213 G x 32021.11 | 77.6 | — | |
| SlowFast R101Pretrain=K400, Params=53.3M2021.12 | 77.6 | — | |
| SlowFast, R101†Pretrain=K400, Frames=16, Params (10^6)=53.32021.10 | 77.6 | — | |
| PEcore GParam.=1.9B, Resolution=256x256, Frame clips=32, Temporal crops=4, Spatial crops=32026.03 | 76.9 | — | |
| Siglip2Type=Discriminative Pretraining, Pretrained on=Images, Arch.=ViT-G, Param.=1.2B, Attentive probing=true2025.12 | 75.3 | — | |
| SigLIP2Param.=1.2B, Resolution=256x256, Frame clips=32, Temporal crops=4, Spatial crops=32026.03 | 75.3 | — | |
| TimeSformerFLOPs x clips=196 G x 32021.11 | 75 | — | |
| TimeSformer†Pretrain=IN-21K, Frames=16, Params (10^6)=1212021.10 | 74.9 | — | |
| MG SamplerFrames=16, Backbone=X3D-M2023.04 | 74.6 | — | |
| UniformFrames=16, Backbone=X3D-M2023.04 | 73.5 | — | |
| K-centeredFrames=16, Backbone=ViT2023.04 | 72.5 | — | |
| VideoPrismType=Discriminative Pretraining, Pretrained on=Videos, Arch.=ViT-G, Param.=1B, Attentive probing=true2025.12 | 71.3 | — | |
| VideoPrismParam.=1B2026.03 | 71.3 | — | |
| RandomFrames=16, Backbone=X3D-M2023.04 | 71.1 | — | |
| MAE-Align w/ real humansPrivacy Preserving=false, Stage 1: MAE=Kinetics, Stage 2: Alignment=Kinetics, Evaluation Protocol=Fine-tuning, Backbone=ViT-B2023.11 | 66.3 | — | |
| Privacy-Preserving MAE-Align (PPMA)Privacy Preserving=true, Stage 1: MAE=NH Kinetics, Stage 2: Alignment=NH Kinetics + Synthetic, Evaluation Protocol=Fine-tuning, Backbone=ViT-B2023.11 | 64 | — | |
| TSN (RN50 backbone)Privacy Preserving=true, Stage 2: Alignment=Synthetic, Evaluation Protocol=Fine-tuning, Backbone=ResNet-502023.11 | 63.5 | — | |
| MAE-Align w/ SyntheticPrivacy Preserving=true, Stage 1: MAE=Synthetic, Stage 2: Alignment=Synthetic, Evaluation Protocol=Fine-tuning, Backbone=ViT-B2023.11 | 61.1 | — | |
| R(2+1)D (RN50 backbone)Privacy Preserving=true, Stage 2: Alignment=Synthetic, Evaluation Protocol=Fine-tuning, Backbone=ResNet-502023.11 | 57.3 | — | |
| I3D (RN50 backbone)Privacy Preserving=true, Stage 2: Alignment=Synthetic, Evaluation Protocol=Fine-tuning, Backbone=ResNet-502023.11 | 55.3 | — | |
| TSN†Pretrain=ImageNet, Frames=32021.10 | 52.5 | — | |
| TSMPretrain=ImageNet, Frames=3, Params (10^6)=42.92021.10 | 51.1 | — | |
| ST-S3D†Pretrain=K400, Frames=82021.10 | 50.6 | — | |
| I3D†Pretrain=K400, Frames=82021.10 | 48.3 | — | |
| TimeSformer KineticsPrivacy Preserving=false, Stage 2: Alignment=Kinetics, Evaluation Protocol=Fine-tuning, Backbone=ViT-B, Pre-trained on ImageNet-21K=true2023.11 | 46.4 | — | |
| TimeSformer SyntheticPrivacy Preserving=false, Stage 2: Alignment=Synthetic, Evaluation Protocol=Fine-tuning, Backbone=ViT-B, Pre-trained on ImageNet-21K=true2023.11 | 44.9 | — | |
| EANPre-train=ImageNet, Frames=16x22021.07 | 41.7 | — | |
| EANPre-train=ImageNet, Frames=162021.07 | 40.4 | — | |
| GSTPre-train=ImageNet, Frames=162021.07 | 38.8 | — | |
| CorrNet-101Frames=32x102021.07 | 38.6 | — | |
| TEA-ResNet50Pre-train=ImageNet, Frames=162021.07 | 36 | — | |
| Kanojia et al.Pre-train=ImageNet, Frames=642021.07 | 35.6 | — | |
| C3DPre-train=ImageNet, Frames=162021.07 | 34.5 | — | |
| P3DPre-train=ImageNet, Frames=162021.07 | 32.4 | — | |
| TIME + V-JEPA 22026.05 | 30.04 | 0.04 | |
| R(2+1)DPre-train=Kinetics2021.07 | 28.9 | — | |
| C3DPre-train=ImageNet, Frames=642021.07 | 27.6 | — | |
| TIME + VideoMAEv22026.05 | 24.97 | 0.73 | |
| TRNPre-train=ImageNet, Frames=82021.07 | 22.8 | — | |
| Privacy-Preserving MAE-Align (PPMA)Privacy Preserving=true, Stage 1: MAE=NH Kinetics, Stage 2: Alignment=NH Kinetics + Synthetic, Evaluation Protocol=Linear Probing, Backbone=ViT-B2023.11 | 21.9 | — | |
| MAE-Align w/ real humansPrivacy Preserving=false, Stage 1: MAE=Kinetics, Stage 2: Alignment=Kinetics, Evaluation Protocol=Linear Probing, Backbone=ViT-B2023.11 | 19.9 | — | |
| TimeSformer SyntheticPrivacy Preserving=false, Stage 2: Alignment=Synthetic, Evaluation Protocol=Linear Probing, Backbone=ViT-B, Pre-trained on ImageNet-21K=true2023.11 | 19.2 | — | |
| TIME + DINOv3frames=4f2026.05 | 18.08 | 1.73 | |
| TimeSformer KineticsPrivacy Preserving=false, Stage 2: Alignment=Kinetics, Evaluation Protocol=Linear Probing, Backbone=ViT-B, Pre-trained on ImageNet-21K=true2023.11 | 17 | — | |
| TSNPre-train=ImageNet, Frames=82021.07 | 16.7 | — | |
| MAE-Align w/ SyntheticPrivacy Preserving=true, Stage 1: MAE=Synthetic, Stage 2: Alignment=Synthetic, Evaluation Protocol=Linear Probing, Backbone=ViT-B2023.11 | 16.7 | — | |
| TIME + CLIPframes=4f2026.05 | 16.11 | 2.14 | |
| TIME + RVM2026.05 | 11.74 | 2.53 | |
| TSN (RN50 backbone)Privacy Preserving=true, Stage 2: Alignment=Synthetic, Evaluation Protocol=Linear Probing, Backbone=ResNet-502023.11 | 10.9 | — | |
| I3D (RN50 backbone)Privacy Preserving=true, Stage 2: Alignment=Synthetic, Evaluation Protocol=Linear Probing, Backbone=ResNet-502023.11 | 10.1 | — | |
| R(2+1)D (RN50 backbone)Privacy Preserving=true, Stage 2: Alignment=Synthetic, Evaluation Protocol=Linear Probing, Backbone=ResNet-502023.11 | 10 | — | |
| AC+HW-JEPAPretraining Datasets=UCF-101 + SSv2 + ImageNet-1002026.05 | 9.39 | 0.71 | |
| HW-JEPAPretraining Datasets=UCF-101 + SSv2 + ImageNet-1002026.05 | 8.98 | 0.3 | |
| BaselinePretraining Datasets=UCF-101 + SSv2 + ImageNet-1002026.05 | 8.68 | — | |
| FWM-HW-LDPretraining Datasets=UCF-101 + SSv2 + ImageNet-1002026.05 | 8.38 | 0.3 |