Video Classification on Something-Something V2 (test)
0.773Top-1 AccMVD-H (Teacher-H)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| MVD-H (Teacher-H)extra data=IN-1K+K400, GFLOPs=1192x6, Param=632, distilled_epochs=8002022.12 | 0.773 | — | |
| MVD-L (Teacher-L)extra data=IN-1K+K400, GFLOPs=597x6, Param=305, distilled_epochs=8002022.12 | 0.767 | — | |
| MVD-L (Teacher-L)extra data=IN-1K+K400, GFLOPs=597x6, Param=3052022.12 | 0.761 | — | |
| OmniMAE ViT-Hextra data=IN-1K, GFLOPs=1192x6, Param=6322022.12 | 0.753 | — | |
| MaskFeat MViT-Lextra data=K400, GFLOPs=2828x3, Param=2182022.12 | 0.744 | — | |
| VideoMAE ViT-Lextra data=None, GFLOPs=597x6, Param=3052022.12 | 0.743 | — | |
| OmniMAE ViT-Lextra data=IN-1K, GFLOPs=597x6, Param=3052022.12 | 0.742 | — | |
| ST-MAE ViT-Hextra data=K400, GFLOPs=1193x3, Param=6322022.12 | 0.741 | — | |
| VideoMAE ViT-Lextra data=K400, GFLOPs=597x6, Param=3052022.12 | 0.74 | — | |
| MVD-B (Teacher-L)extra data=IN-1K+K400, GFLOPs=180x6, Param=872022.12 | 0.737 | — | |
| TAdaFormer-L/14#frames=32, GFLOPs x views=1716x3x2, Initialization=CLIP-400M2023.08 | 0.736 | — | |
| MViTv2-Lpretrain=IN21K + K400, Resolution=312x312, Frame config=40x3, FLOPs x views=2828x3x1, Param=213.12021.12 | 0.733 | 0.941 | |
| MViTv2-Lextra data=IN-21K+K400, GFLOPs=2828x3, Param=2132022.12 | 0.733 | — | |
| MViTv2-L#frames=40, GFLOPs x views=2828x3x1, Resolution=312x3122023.08 | 0.733 | — | |
| UniFormerV2-L/14#frames=32, GFLOPs x views=~1716x3x2, Initialization=CLIP-400M2023.08 | 0.731 | — | |
| MVD-B (Teacher-B)extra data=IN-1K+K400, GFLOPs=180x6, Param=872022.12 | 0.725 | — | |
| TAdaFormer-L/14#frames=16, GFLOPs x views=858x3x2, Initialization=CLIP-400M2023.08 | 0.724 | — | |
| ST-Adapter-L/14#frames=32, GFLOPs x views=2749x3x1, Initialization=CLIP-400M2023.08 | 0.723 | — | |
| MViTv2-B, 32x3pretrain=IN21K + K400, FLOPs x views=225x3x1, Param=51.12021.12 | 0.721 | 0.934 | |
| ST-MAE ViT-Lextra data=K400, GFLOPs=598x3, Param=3042022.12 | 0.721 | — | |
| BEVT Swin-Bextra data=IN-1K+K400, GFLOPs=321x3, Param=882022.12 | 0.714 | — | |
| UniFormer-BPre-training=K400, Sampling (frame x crop x clip)=32x3x2, FLOPs (G)=15542022.01 | 0.714 | 0.928 | |
| TAdaFormer-B/16#frames=32, GFLOPs x views=374x3x2, Initialization=CLIP-400M2023.08 | 0.713 | — | |
| UniFormer-BPre-training=K600, Sampling (frame x crop x clip)=32x3x2, FLOPs (G)=15542022.01 | 0.713 | 0.928 | |
| Uniformer-Bextra data=K400, GFLOPs=259x3, Param=502022.12 | 0.712 | — | |
| UniFormer-BPre-training=K400, Sampling (frame x crop x clip)=32x3x1, FLOPs (G)=7772022.01 | 0.712 | 0.928 | |
| UniFormer-BPre-training=K600, Sampling (frame x crop x clip)=32x3x1, FLOPs (G)=7772022.01 | 0.712 | 0.928 | |
| TAdaConvNeXtV2-B#frames=32, GFLOPs x views=324x3x2, Initialization=ImageNet21K+K4002023.08 | 0.711 | — | |
| UniFormerV2-B/16#frames=32, GFLOPs x views=~370x3x2, Initialization=CLIP-400M2023.08 | 0.71 | — | |
| MVD-S (Teacher-L)extra data=IN-1K+K400, GFLOPs=57x6, Param=222022.12 | 0.709 | — | |
| VideoMAE ViT-Bextra data=None, GFLOPs=180x6, Param=872022.12 | 0.708 | — | |
| MVD-S (Teacher-B)extra data=IN-1K+K400, GFLOPs=57x6, Param=222022.12 | 0.707 | — | |
| TAdaConvNeXtV2-S#frames=32, GFLOPs x views=183x3x2, note=second entry in table2023.08 | 0.706 | — | |
| MViTv2-B, 32x3FLOPs x views=225x3x1, Param=51.12021.12 | 0.705 | 0.927 | |
| MViTv2-Bextra data=K400, GFLOPs=225x3, Param=512022.12 | 0.705 | — | |
| MViTv2-B#frames=32, GFLOPs x views=225x3x12023.08 | 0.705 | — | |
| TAdaFormer-B/16#frames=16, GFLOPs x views=187x3x2, Initialization=CLIP-400M2023.08 | 0.704 | — | |
| UniFormer-BPre-training=K400, Sampling (frame x crop x clip)=16x3x1, FLOPs (G)=2902022.01 | 0.704 | 0.928 | |
| ILA-ViT-L/14@336pxPretrain=CLIP-400M, Frames=16, Views=4x3, FLOPs (G)=37232023.04 | 0.702 | 0.918 | |
| UniFormer-BPre-training=K600, Sampling (frame x crop x clip)=16x3x1, FLOPs (G)=2902022.01 | 0.702 | 0.93 | |
| TAdaConvNeXtV2-S#frames=32, GFLOPs x views=183x3x22023.08 | 0.7 | — | |
| TAdaConvNeXtV2-T#frames=32, GFLOPs x views=94x3x22023.08 | 0.698 | — | |
| VideoMAE ViT-Bextra data=K400, GFLOPs=180x6, Param=872022.12 | 0.697 | — | |
| Swin-Bpretrain=IN21K + K400, FLOPs x views=321x3x1, Param=88.82021.12 | 0.696 | 0.927 | |
| Video Swin-BPre-train=K-400, GFLOPs=321, Views=1 x 3, Params=88.82022.06 | 0.696 | 0.927 | |
| TDN R101extra data=IN-1K, GFLOPs=198x3, Param=882022.12 | 0.696 | — | |
| VideoSwin-Bextra data=IN-21K+K400, GFLOPs=321x3, Param=882022.12 | 0.696 | — | |
| Swin-B#frames=32, GFLOPs x views=321x3x12023.08 | 0.696 | — | |
| Swin-BPre-training=K400, Sampling (frame x crop x clip)=32x3x1, FLOPs (G)=9632022.01 | 0.696 | 0.927 | |
| Swin-BPre-trained=Kinetics-4002026.01 | 0.696 | — | |
| OmniMAE ViT-Bextra data=IN-1K, GFLOPs=180x6, Param=872022.12 | 0.695 | — | |
| ST-Adapter-B/16#frames=32, GFLOPs x views=651x3x1, Initialization=CLIP-400M2023.08 | 0.695 | — | |
| AIM-ViT-L/14Pretrain=CLIP-400M, Frames=32, Views=1x3, FLOPs (G)=38362023.04 | 0.694 | 0.923 | |
| UniFormer-SPre-training=K600, Sampling (frame x crop x clip)=16x3x1, FLOPs (G)=1252022.01 | 0.694 | 0.921 | |
| OmniMAE ViT-Bextra data=IN-1K+K400, GFLOPs=180x6, Param=872022.12 | 0.69 | — | |
| MViTv1-B-24, 32x3pretrain=K600, FLOPs x views=236.0x3x1, Param=53.22021.12 | 0.687 | 0.915 | |
| MViT-B-24, 32 x 3Pre-train=K-600, GFLOPs=236, Views=1 x 3, Params=53.22022.06 | 0.687 | 0.915 | |
| MViT-B-24, 32x3Pre-training=K600, Sampling (frame x crop x clip)=32x1x3, FLOPs (G)=7082022.01 | 0.687 | 0.915 | |
| MLP-3D-LPre-train=IN-1K, GFLOPs=336, Views=1 x 3, Params=149.42022.06 | 0.685 | 0.92 | |
| MTV-B#frames=32, GFLOPs x views=930x3x4, Resolution=320x3202023.08 | 0.685 | — | |
| TAdaConvNeXtV2-S#frames=16, GFLOPs x views=91x3x22023.08 | 0.684 | — | |
| MViTv2-S, 16x4FLOPs x views=64.5x3x1, Param=34.42021.12 | 0.682 | 0.914 | |
| TDN-R101#frames=8+16, GFLOPs x views=258x1x12023.08 | 0.682 | — | |
| TDNENPre-training=IN-1K, Sampling (frame x crop x clip)=8+16, FLOPs (G)=1982022.01 | 0.682 | 0.916 | |
| MViTv2-SPre-trained=Kinetics-400, Input configuration=16x42026.01 | 0.682 | — | |
| Mformer-Lextra data=IN-21K+K400, GFLOPs=1185x3, Param=3822022.12 | 0.681 | — | |
| VIMPAC ViT-Lextra data=HowTo100M, GFLOPs=N/Ax30, Param=3072022.12 | 0.681 | — | |
| Mformer-LPretrain=IN-21K+K400, Frames=32, Views=1x3, FLOPs (G)=11852023.04 | 0.681 | 0.912 | |
| Mformer-LPre-training=K400, Sampling (frame x crop x clip)=32x3x1, FLOPs (G)=35552022.01 | 0.681 | 0.912 | |
| MLP-3D-MPre-train=IN-1K, GFLOPs=183, Views=1 x 3, Params=88.32022.06 | 0.68 | 0.917 | |
| EVL-ViT-L/14@336pxPretrain=CLIP-400M, Frames=32, Views=1x3, FLOPs (G)=80902023.04 | 0.68 | — | |
| ILA-ViT-L/14Pretrain=CLIP-400M, Frames=8, Views=4x3, FLOPs (G)=9072023.04 | 0.678 | 0.905 | |
| CT-NetENPre-training=IN-1K, Sampling (frame x crop x clip)=8+12+16+24, FLOPs (G)=2802022.01 | 0.678 | 0.911 | |
| MVIT-BClip Size=64 x 224², Additional Data (# Samples)=K400 (240K), Pretraining=supervised2021.06 | 0.677 | 0.909 | |
| MViTv1-B, 64x3pretrain=K400, FLOPs x views=454x3x1, Param=36.62021.12 | 0.677 | 0.909 | |
| MViT-B, 64 x 3Pre-train=K-400, GFLOPs=455, Views=1 x 3, Params=36.62022.06 | 0.677 | 0.909 | |
| MViTv1-Bextra data=None, GFLOPs=455x3, Param=372022.12 | 0.677 | — | |
| MViT-B, 64x3Pre-training=K400, Sampling (frame x crop x clip)=64x1x3, FLOPs (G)=13652022.01 | 0.677 | 0.909 | |
| UniFormer-SPre-training=K400, Sampling (frame x crop x clip)=16x3x1, FLOPs (G)=1252022.01 | 0.677 | 0.914 | |
| MLP-3D-SPre-train=IN-1K, GFLOPs=108, Views=1 x 3, Params=74.12022.06 | 0.672 | 0.913 | |
| TAdaConvNeXtV2-T#frames=16, GFLOPs x views=47x3x22023.08 | 0.672 | — | |
| X-ViT#frames=16, GFLOPs x views=283x3x12023.08 | 0.672 | — | |
| TAdaConvNeXt-T#frames=32, GFLOPs x views=94x3x22023.08 | 0.671 | — | |
| ATAPretrain=IN-21K, Frames=32, Views=4x3, FLOPs (G)=7932023.04 | 0.671 | 0.908 | |
| MViTv1-BPretrain=K400, Frames=32, Views=1x3, FLOPs (G)=1702023.04 | 0.671 | 0.908 | |
| Mformer-HRPretrain=IN-21K+K400, Frames=64, Views=1x3, FLOPs (G)=9592023.04 | 0.671 | 0.906 | |
| Mformer-HRPre-training=K400, Sampling (frame x crop x clip)=16x3x1, FLOPs (G)=28762022.01 | 0.671 | 0.906 | |
| bLResNetPre-trained=SSv2, Input configuration=32x22026.01 | 0.671 | — | |
| ST Swin w/ LSTCLClip Size=16 x 224², Additional Data (# Samples)=K400 (240K), Pretraining=unsupervised2021.06 | 0.67 | 0.905 | |
| TDNPre-train=IN-1K, GFLOPs=132, Views=1 x 12022.06 | 0.669 | 0.909 | |
| VideoMAE ViT-Sextra data=None, GFLOPs=57x6, Param=222022.12 | 0.668 | — | |
| ILA-ViT-B/16Pretrain=CLIP-400M, Frames=16, Views=4x3, FLOPs (G)=4382023.04 | 0.668 | 0.903 | |
| EVL-ViT-L/14Pretrain=CLIP-400M, Frames=32, Views=1x3, FLOPs (G)=32162023.04 | 0.667 | — | |
| MformerClip Size=16 x 224², Additional Data (# Samples)=ImageNet-21K + K400 (14.2M), Pretraining=supervised2021.06 | 0.665 | 0.901 | |
| Mformer-Bextra data=IN-21K+K400, GFLOPs=370x3, Param=1092022.12 | 0.665 | — | |
| Mformer-BPretrain=IN-21K+K400, Frames=16, Views=1x3, FLOPs (G)=3702023.04 | 0.665 | 0.901 | |
| VideoMAE ViT-Sextra data=K400, GFLOPs=57x6, Param=222022.12 | 0.664 | — | |
| AIM-ViT-B/16Pretrain=CLIP-400M, Frames=8, Views=1x3, FLOPs (G)=2082023.04 | 0.664 | 0.905 | |
| MLP-3D-XSPre-train=IN-1K, GFLOPs=60, Views=1 x 3, Params=55.12022.06 | 0.66 | 0.904 | |
| TAM R50extra data=IN-1K, GFLOPs=99x6, Param=512022.12 | 0.66 | — |