Action Recognition on Something-Something v2 (val)
79.8Top-1 AccuracyFTP-UniFormerV2-L/14
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| FTP-UniFormerV2-L/14Input=(32 x 224^2), Crops=1x3, Params (M)=620, TFLOPs=6.62024.03 | 79.8 | 98.9 | |
| FTP-UniFormerV2-B/16Input=(16 x 224^2), Crops=1x3, Params (M)=183, TFLOPs=1.12024.03 | 77.3 | 96.7 | |
| InternVideoInput=(64 x 224^2), Crops=16x4, Params (M)=1300, TFLOPs=86.22024.03 | 77.2 | 95.9 | |
| VideoMAE V2-gInput=(64 x 266^2), Crops=5x3, Params (M)=1050, TFLOPs=160.32024.03 | 77 | 95.9 | |
| TubeViT-LInput=(32 x 224^2), Crops=4x3, Params (M)=311, TFLOPs=9.52024.03 | 76.1 | 95.2 | |
| FluxViT-BExtra Data=K710+MASH, GFLOPS=440×6, more informative tokens=true2025.03 | 75.6 | 95.1 | |
| MAEpre-train=MAE, pre-train data=K700, architecture=ViT-H, input size=16 x 224^2, FLOPs=1193 x 3 x 1, param.=6322022.05 | 75.5 | 95 | |
| MAEpre-train=MAE, pre-train data=K700, architecture=ViT-H, input size=16 x 224^2, FLOPs=1193 x 3 x 1, param.=6322022.05 | 75.5 | 95 | |
| FluxViT-BExtra Data=K710+MASH, GFLOPS=255×6, more informative tokens=true2025.03 | 75.5 | 95.1 | |
| VideoMAEBackbone=ViT-L, Extra data=no external data, Ex. labels=X, Frames=32, GFLOPs=1436×1×3, Param=3052022.03 | 75.4 | 95.2 | |
| VideoMAE-LInput=(32 x 224^2), Crops=1x3, Params (M)=305, TFLOPs=4.32024.03 | 75.4 | 95.2 | |
| FluxViT-BExtra Data=K710+MASH, GFLOPS=440×6, more informative tokens=false2025.03 | 75.3 | 95.1 | |
| VJEPA2 ViT-gEncoder Params=1B, Pre-train data=VM22M, Frozen evaluation=true, Attentive probe=true2026.06 | 75.3 | — | |
| MAEpre-train=MAE, pre-train data=K600, architecture=ViT-H, input size=16 x 224^2, FLOPs=1193 x 3 x 1, param.=6322022.05 | 75.2 | 94.9 | |
| MAE (ViT-H)pre-train=MAE, pre-train data=K600, architecture=ViT-H, input size=16 x 224^2, FLOPs=1193 x 3 x 1, param.=6322022.05 | 75.2 | 94.9 | |
| Side4VideoEvaluation Protocol=Frozen backbone, Backbones=ViT-E/14, Pre-train=LAION-2B, Frames × Views=16×3×2, TFLOPS=15.96×62023.11 | 75.2 | 94 | |
| FluxViT-BExtra Data=K710+MASH, GFLOPS=255×6, more informative tokens=false2025.03 | 75.1 | 94.9 | |
| FluxViT-BExtra Data=K710+MASH, GFLOPS=108×6, more informative tokens=true2025.03 | 75.1 | 94.8 | |
| MaskFeat↑312Backbone=MVIT-L, Extra data=Kinetics-600, Ex. labels=✓, Frames=40, GFLOPs=2828×1×3, Param=2182022.03 | 75 | 95 | |
| MaskFeatpre-train=MaskFeat, pre-train data=K600, architecture=MViTv2-L, input size=40 x 312^2, FLOPs=2828 x 3 x 1, param.=2182022.05 | 75 | 95 | |
| MaskFeat (MViTv2-L)pre-train=MaskFeat, pre-train data=K600, architecture=MViTv2-L, input size=40 x 312^2, FLOPs=2828 x 3 x 1, param.=2182022.05 | 75 | 95 | |
| MaskFeat (MViT-L↑312, 40x3)pre-train=MaskFeat, K600, FLOPs=2828, Param=2182021.12 | 75 | 95 | |
| MaskFeat-LInput=(64 x 312^2), Crops=4x3, Params (M)=218, TFLOPs=8.52024.03 | 75 | 95 | |
| ATMEvaluation Protocol=Full Fine-tuning, Backbones=ViT-L/14, Pre-train=Merged-2B, Frames × Views=16×3×2, TFLOPS=0.84×62023.11 | 74.6 | 94.4 | |
| MaskFeatpre-train=MaskFeat, pre-train data=K400, architecture=MViTv2-L, input size=40 x 312^2, FLOPs=2828 x 3 x 1, param.=2182022.05 | 74.4 | 94.6 | |
| MaskFeat (MViT-L↑312, 40x3)pre-train=MaskFeat, K400, FLOPs=2828, Param=2182021.12 | 74.4 | 94.6 | |
| VideoMAEBackbone=ViT-L, Extra data=no external data, Ex. labels=X, Frames=16, GFLOPs=597×2×3, Param=3052022.03 | 74.3 | 94.6 | |
| Side4VideoEvaluation Protocol=Frozen backbone, Backbones=ViT-E/14, Pre-train=LAION-2B, Frames × Views=8×3×2, TFLOPS=7.98×62023.11 | 74.3 | 94 | |
| MAEpre-train=MAE, pre-train data=K400, architecture=ViT-H, input size=16 x 224^2, FLOPs=1193 x 3 x 1, param.=6322022.05 | 74.1 | 94.5 | |
| TDS-CLIP-L/14Training Protocol=Frozen CLIP, Pre-train=CLIP-400M, Frames × Cr. × Cl.=16×3×2, TFLOPS=1.6×62024.08 | 74.1 | 94 | |
| VideoMAEBackbone=ViT-L, Extra data=Kinetics-400, Ex. labels=X, Frames=16, GFLOPs=597×2×3, Param=3052022.03 | 74 | 94.6 | |
| VideoMAE-LExtra Data=K400, GFLOPS=596×62025.03 | 74 | 94.6 | |
| FluxViT-BExtra Data=K710+MASH, GFLOPS=49×6, more informative tokens=true2025.03 | 73.9 | 94.4 | |
| MJEPA ViT-g + data scalingEncoder Params=1B, Pre-train data=AS+VM2M, Frozen evaluation=true, Attentive probe=true2026.06 | 73.9 | — | |
| FluxViT-SExtra Data=K710+MASH, GFLOPS=154×6, more informative tokens=true2025.03 | 73.8 | 94.1 | |
| VJEPA2 ViT-LEncoder Params=300M, Pre-train data=VM22M, Frozen evaluation=true, Attentive probe=true2026.06 | 73.7 | — | |
| MAEpre-train=MAE, pre-train data=K700, architecture=ViT-L, input size=16 x 224^2, FLOPs=598 x 3 x 1, param.=3042022.05 | 73.6 | 94.4 | |
| MAEpre-train=MAE, pre-train data=K700, architecture=ViT-L, input size=16 x 224^2, FLOPs=598 x 3 x 1, param.=3042022.05 | 73.6 | 94.4 | |
| ATMEvaluation Protocol=Full Fine-tuning, Backbones=ViT-L/14, Pre-train=CLIP-400M, Frames × Views=16×3×2, TFLOPS=0.84×62023.11 | 73.5 | 93.7 | |
| InternVideo2-BExtra Data=K710+MASH, GFLOPS=253×62025.03 | 73.5 | 94.4 | |
| FluxViT-SExtra Data=K710+MASH, GFLOPS=154×6, more informative tokens=false2025.03 | 73.4 | 94.1 | |
| FluxViT-SExtra Data=K710+MASH, GFLOPS=83×6, more informative tokens=true2025.03 | 73.4 | 94.1 | |
| TDS-CLIP-L/14Training Protocol=Frozen CLIP, Pre-train=CLIP-400M, Frames × Cr. × Cl.=8×3×2, TFLOPS=0.8×62024.08 | 73.4 | 93.8 | |
| MViTv2-Lpre-train=supervised, pre-train data=K400 + IN21K, architecture=MViTv2-L, input size=40 x 224^2, FLOPs=2828 x 3 x 1, param.=2132022.05 | 73.3 | 94.1 | |
| MViT-L↑312, 40x3pre-train=Sup., IN-21K+K400, FLOPs=2828, Param=2182021.12 | 73.3 | 94.1 | |
| MJEPA ViT-L + data scalingEncoder Params=300M, Pre-train data=AS+VM2M, Frozen evaluation=true, Attentive probe=true2026.06 | 73.3 | — | |
| Side4VideoEvaluation Protocol=Frozen backbone, Backbones=ViT-L/14, Pre-train=CLIP-400M, Frames × Views=16×3×2, TFLOPS=1.74×62023.11 | 73.2 | 93.9 | |
| DiSTEvaluation Protocol=Frozen backbone, Backbones=ViT-L/14, Pre-train=CLIP-400M, Frames × Views=32×1×3, TFLOPS=2.83×32023.11 | 73.1 | 93.2 | |
| MAEpre-train=MAE, pre-train data=K600, architecture=ViT-L, input size=16 x 224^2, FLOPs=598 x 3 x 1, param.=3042022.05 | 73 | 94.2 | |
| MAE (ViT-L)pre-train=MAE, pre-train data=K600, architecture=ViT-L, input size=16 x 224^2, FLOPs=598 x 3 x 1, param.=3042022.05 | 73 | 94.2 | |
| UniFormerV2Evaluation Protocol=Full Fine-tuning, Backbones=ViT-L/14, Pre-train=CLIP-400M, Frames × Views=32×3×1, TFLOPS=1.73×32023.11 | 73 | 94.5 | |
| UniFormerV2-L/14Input=(32 x 224^2), Crops=1x3, Params (M)=574, TFLOPs=5.22024.03 | 73 | 94.5 | |
| UniFormerV2-LExtra Data=CLIP-400M, GFLOPS=1718×32025.03 | 73 | 94.5 | |
| FluxViT-SExtra Data=K710+MASH, GFLOPS=83×6, more informative tokens=false2025.03 | 72.9 | 94 | |
| FluxViT-SExtra Data=K710+MASH, GFLOPS=32×6, more informative tokens=true2025.03 | 72.5 | 93.8 | |
| SMILEBackbone=ViT-B, Epochs=800, Pretrain=SSv2, Params=87M2025.04 | 72.5 | — | |
| MVDBackbone=ViT-B, Epochs=1600+400, Pretrain=K400, Params=87M2025.04 | 72.5 | — | |
| SMILEBackbone=ViT-B, Epochs=800, Pretrain=SSv22025.04 | 72.5 | — | |
| DiST-L/14Training Protocol=Frozen CLIP, Pre-train=CLIP-400M, Frames × Cr. × Cl.=16×3×1, TFLOPS=1.4×32024.08 | 72.5 | 93 | |
| SMILEBackbone=ViT-B, Epochs=1200, Pretrain=K400, Params=87M2025.04 | 72.4 | — | |
| ViT-L w/ ST-AdapterPretrain=CLIP, #Frames=32×3×1, GFlops=8248, Evaluation Protocol=frozen backbone2022.06 | 72.3 | 93.9 | |
| ST-AdapterEvaluation Protocol=Frozen backbone, Backbones=ViT-L/14, Pre-train=CLIP-400M, Frames × Views=32×1×3, TFLOPS=2.75×32023.11 | 72.3 | 93.9 | |
| MGMAEBackbone=ViT-B, Epochs=2400, Pretrain=SSv2, Params=87M2025.04 | 72.3 | — | |
| MViTv2-Bpre-train=supervised, pre-train data=K400 + IN21K, architecture=MViTv2-B, input size=32 x 224^2, FLOPs=225 x 3 x 1, param.=512022.05 | 72.1 | 93.4 | |
| MAEpre-train=MAE, pre-train data=K400, architecture=ViT-L, input size=16 x 224^2, FLOPs=598 x 3 x 1, param.=3042022.05 | 72.1 | 93.9 | |
| MViTv2-BInput=(32 x 224^2), Crops=1x3, Params (M)=51, TFLOPs=0.72024.03 | 72.1 | 93.4 | |
| MGMBackbone=ViT-B, Epochs=2400, Pretrain=SSv2, Params=87M2025.04 | 72.1 | — | |
| SMILEBackbone=ViT-B, Epochs=600, Pretrain=K400, Params=87M2025.04 | 72.1 | — | |
| SMILEBackbone=ViT-B, Epochs=600, Pretrain=K4002025.04 | 72.1 | — | |
| UniFormer V2-L/14Training Protocol=Full Fine-tuning, Pre-train=CLIP-400M, Frames × Cr. × Cl.=16×3×1, TFLOPS=0.9×32024.08 | 72.1 | 93.6 | |
| TDS-CLIP-B/16Training Protocol=Frozen CLIP, Pre-train=CLIP-400M, Frames × Cr. × Cl.=16×3×2, TFLOPS=0.4×62024.08 | 72.1 | 93.3 | |
| Uniformerv2Backbone=ViT-L/14, Giga FLOP=2600, Train. Par. (M)=574, Views=16*1*32026.02 | 72.1 | 93.6 | |
| Frame2Freq-MSBackbone=ViT-L/14, Giga FLOP=1322, Train. Par. (M)=19, Views=16*1*32026.02 | 72.1 | 93.2 | |
| FluxViT-BExtra Data=K710+MASH, GFLOPS=108×6, more informative tokens=false2025.03 | 72 | 93.3 | |
| ViT-L w/ ST-AdapterPretrain=CLIP, #Frames=16×3×1, GFlops=4124, Evaluation Protocol=frozen backbone2022.06 | 71.9 | 93.4 | |
| Side4VideoEvaluation Protocol=Frozen backbone, Backbones=ViT-L/14, Pre-train=CLIP-400M, Frames × Views=8×3×2, TFLOPS=0.87×62023.11 | 71.9 | 93.5 | |
| ST-Adapter-L/14Training Protocol=Frozen CLIP, Pre-train=CLIP-400M, Frames × Cr. × Cl.=16×3×1, TFLOPS=1.4×32024.08 | 71.9 | 93.4 | |
| TDS-CLIP-B/16Training Protocol=Frozen CLIP, Pre-train=CLIP-400M, Frames × Cr. × Cl.=8×3×2, TFLOPS=0.2×62024.08 | 71.8 | 93 | |
| Side4VideoEvaluation Protocol=Frozen backbone, Backbones=ViT-B/16, Pre-train=CLIP-400M, Frames × Views=16×3×2, TFLOPS=0.36×62023.11 | 71.5 | 92.8 | |
| InternVideo2-SExtra Data=K710+MASH, GFLOPS=83×62025.03 | 71.5 | 93.4 | |
| Side4Video-B/16Training Protocol=Frozen CLIP, Pre-train=CLIP-400M, Frames × Cr. × Cl.=16×3×2, TFLOPS=0.4×62024.08 | 71.5 | 92.8 | |
| BEVTPretrain=IN-1K+K400, GFLOPS=321, Crops=3, tokenizer=PeCo [18]2021.12 | 71.4 | — | |
| OMNIVORE (Swin-B)Pretrain=IN22K+K400+SUN, #Frames=32×3×1, GFlops=963, Evaluation Protocol=full-finetuning2022.06 | 71.4 | 93.5 | |
| BEVTpre-train=BEVT, pre-train data=K400 + IN1K, architecture=Swin-B, input size=32 x 224^2, FLOPs=321 x 3 x 1, param.=882022.05 | 71.4 | — | |
| BEVTInput=(32 x 224^2), Params (M)=88, TFLOPs=1.02024.03 | 71.4 | — | |
| VJEPA ViT-HEncoder Params=600M, Pre-train data=VM2M, Frozen evaluation=true, Attentive probe=true2026.06 | 71.4 | — | |
| UniFormer-BPretrain=IN1K+K600, #Frames=32×3×1, GFlops=777, Evaluation Protocol=full-finetuning2022.06 | 71.2 | 92.8 | |
| UniFormer V1-BInput=(32 x 224^2), Crops=1x3, Params (M)=50, TFLOPs=0.82024.03 | 71.2 | 92.8 | |
| VideoMAE-v2-B/16Frames=16, Views=2 x 32024.03 | 71.2 | — | |
| Uniformer-BBackbone=Uformer-B, Pretrain=K400, Params=50M2025.04 | 71.2 | — | |
| SIGMABackbone=ViT-B, Epochs=800, Pretrain=SSv2, Params=87M2025.04 | 71.2 | — | |
| SIGMABackbone=ViT-B, Epochs=800, Pretrain=SSv22025.04 | 71.2 | — | |
| SIGMABackbone=ViT-B, Epochs=800, Pretrain=K400, Params=87M2025.04 | 71.1 | — | |
| MGMBackbone=ViT-B, Epochs=800, Pretrain=K400, Params=87M, Note=re-evaluated by current paper2025.04 | 71.1 | — | |
| SIGMABackbone=ViT-B, Epochs=800, Pretrain=K4002025.04 | 71.1 | — | |
| MGMBackbone=ViT-B, Epochs=800, Pretrain=K400, evaluated_by_authors=true2025.04 | 71.1 | — | |
| MGMAEBackbone=ViT-B, Epochs=800, Pretrain=SSv2, Params=87M2025.04 | 71 | — | |
| MGMAEBackbone=ViT-B, Epochs=800, Pretrain=SSv22025.04 | 71 | — | |
| FluxViT-SExtra Data=K710+MASH, GFLOPS=13×6, more informative tokens=true2025.03 | 70.9 | 93.1 | |
| VideoMAEBackbone=ViT-B, Extra data=no external data, Ex. labels=X, Frames=16, GFLOPs=180×2×3, Param=872022.03 | 70.8 | 92.4 |