Video Classification on Kinetics-600 (test)
86.3Top-1 AccuracyTL 16at18 w. Fuser (L/10)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| TL 16at18 w. Fuser (L/10)GFLOPS=4100, number of views=12, Tokens=16, Insertion Layer=18, Backbone=ViT-L/10, TokenFuser=true2021.06 | 86.3 | 97 | — | |
| Swin-L (384)GFLOPS=2107, number of views=502021.06 | 86.1 | 97.3 | — | |
| TL 16at18 (L/10)GFLOPS=4076, number of views=12, Tokens=16, Insertion Layer=18, Backbone=ViT-L/102021.06 | 86.1 | 97 | — | |
| TL 8at18 (L/16)GFLOPS=1105, number of views=12, Tokens=8, Insertion Layer=18, Backbone=ViT-L/162021.06 | 86 | 97 | — | |
| Swin-L (384)GFLOPS=2107, number of views=122021.06 | 85.9 | 97.1 | — | |
| TL 16at12 (L/16)GFLOPS=766, number of views=12, Tokens=16, Insertion Layer=12, Backbone=ViT-L/162021.06 | 84.4 | 96 | — | |
| ViViT-L/16GFLOPS=1446, number of views=122021.06 | 84.3 | 96.2 | — | |
| Swin-BGFLOPS=282, number of views=122021.06 | 84 | 96.5 | — | |
| MVIT-B-24, 32x3Mem Max (GB)=4.40, Max BS=2, GFLOPs x views=236x1x5, Param (M)=52.92023.02 | 83.8 | — | — | |
| Rev-MViT-B-24, 32x3Mem Max (GB)=1.64, Max BS=7, GFLOPs x views=223x1x5, Param (M)=51.82023.02 | 83.7 | — | — | |
| ST Swin w/ LSTCLClip Size=16 x 224^22021.06 | 83.6 | 96.6 | — | |
| VATT-LClip Size=32 x 320^2, Additional Data=AudioSet + HowTo100M (3.2M)2021.06 | 83.6 | 96.6 | — | |
| MVIT-BClip Size=32 x 224^22021.06 | 83.4 | 96.3 | — | |
| MVIT-B-16, 32x3GFLOPs x views=170x1x5, Param (M)=36.82023.02 | 83.4 | — | — | |
| ViT-L-ViViT-IN-21KGFLOPs x views=3992x3x4, Param (M)=310.82023.02 | 83 | — | — | |
| ViViT-LClip Size=16 x 224^2, Additional Data=ImageNet-21K (14M)2021.06 | 82.5 | 95.6 | — | |
| TimeSformer-HRGFLOPS=1703, number of views=32021.06 | 82.4 | 96 | — | |
| ViT-B-TimeSformer-IN-21KGFLOPs x views=1703x3x1, Param (M)=121.42023.02 | 82.4 | — | — | |
| TimeSformer-LClip Size=96 x 224^2, Additional Data=ImageNet-21K (14M)2021.06 | 82.2 | 95.6 | — | |
| MVIT-BClip Size=16 x 224^22021.06 | 82.1 | 95.7 | — | |
| ST Swin w/ LSTCLClip Size=8 x 224^22021.06 | 82 | 95.5 | — | |
| X3D-XLClip Size=16 x 312^22021.06 | 81.9 | 95.9 | — | |
| X3D-XLGFLOPS=48, number of views=302021.06 | 81.9 | 95.5 | — | |
| X3D-XLGFLOPs x views=48.4x3x10, Param (M)=11.02023.02 | 81.9 | — | — | |
| SlowFastClip Size=16 x 224^22021.06 | 81.8 | 95.1 | — | |
| SlowFast 16x8, R101+NLGFLOPS=234, number of views=302021.06 | 81.8 | 95.1 | — | |
| SlowFast 16x8+NLGFLOPs x views=234x3x10, Param (M)=59.92023.02 | 81.8 | — | — | |
| MformerClip Size=16 x 224^2, Additional Data=ImageNet-21K (14M)2021.06 | 81.6 | 95.6 | — | |
| MVIT-B-16, 16x4GFLOPs x views=70.3x1x5, Param (M)=36.62023.02 | 81.3 | — | — | |
| VATT-BClip Size=32 x 320^2, Additional Data=AudioSet + HowTo100M (3.2M)2021.06 | 80.5 | 95.5 | — | |
| SlowFastClip Size=8 x 224^22021.06 | 80.4 | 94.8 | — | |
| TimeSformerClip Size=8 x 224^2, Additional Data=ImageNet-21K (14M)2021.06 | 79.1 | 94.4 | — | |
| ST Swin from scratchClip Size=8 x 224^22021.06 | 74.7 | 92.2 | — | |
| EVA-02-CLIPZero-shot=true, Backbone=ViT-E/14+, Training Scale=Enlarged2023.03 | 69.3 | — | — | |
| EVA-02-CLIPZero-shot=true, Backbone=ViT-E/14, Training Scale=Enlarged2023.03 | 68.6 | — | — | |
| EVA-01-CLIPZero-shot=true, Backbone=ViT-g/14+, Training Scale=Enlarged2023.03 | 67 | — | — | |
| EVA-02-CLIPZero-shot=true, Backbone=ViT-L/14+, Training Scale=Enlarged2023.03 | 66.1 | — | — | |
| Open CLIPZero-shot=true, Backbone=ViT-G/14, Training Scale=Enlarged2023.03 | 66.1 | — | — | |
| OpenAI CLIPZero-shot=true, Backbone=ViT-L/14+, Training Scale=Enlarged2023.03 | 65 | — | — | |
| EVA-02-CLIPZero-shot=true, Backbone=ViT-L/14, Training Scale=Standard2023.03 | 64.9 | — | — | |
| EVA-01-CLIPZero-shot=true, Backbone=ViT-g/14, Training Scale=Enlarged2023.03 | 64.4 | — | — | |
| OpenAI CLIPZero-shot=true, Backbone=ViT-L/14, Training Scale=Standard2023.03 | 64.2 | — | — | |
| Open CLIPZero-shot=true, Backbone=ViT-g/14, Training Scale=Enlarged2023.03 | 64.1 | — | — | |
| Open CLIPZero-shot=true, Backbone=ViT-H/14, Training Scale=Enlarged2023.03 | 63.6 | — | — | |
| Open CLIPZero-shot=true, Backbone=ViT-L/14, Training Scale=Standard2023.03 | 58.6 | — | — | |
| EVA-02-CLIPZero-shot=true, Backbone=ViT-B/16, Training Scale=Standard2023.03 | 57 | — | — | |
| OpenAI CLIPZero-shot=true, Backbone=ViT-B/16, Training Scale=Standard2023.03 | 56.5 | — | — | |
| Open CLIPZero-shot=true, Backbone=ViT-B/16, Training Scale=Standard2023.03 | 54.2 | — | — | |
| BIKEEncoder=ViT-L/14, Zero-shot protocol=ZSVR2023.08 | — | — | 67 | |
| ERZero-shot=true2022.07 | — | — | 42.1 | |
| EREncoder=TSM, Zero-shot protocol=ZSVR2023.08 | — | — | 42.1 | |
| OTIEncoder=ViT-B/32, Zero-shot protocol=ZSVR2023.08 | — | — | 64.5 | |
| OTIEncoder=ViT-B/16, Zero-shot protocol=ZSVR2023.08 | — | — | 66.9 | |
| OTIEncoder=ViT-L/14, Zero-shot protocol=ZSVR2023.08 | — | — | 70.6 | |
| Text4VisZero-shot=true2022.07 | — | — | 68.9 | |
| Text4VisEncoder=ViT-L/14, Zero-shot protocol=ZSVR2023.08 | — | — | 68.9 | |
| X-CLIPEncoder=ViT-B/16, Zero-shot protocol=ZSVR2023.08 | — | — | 65.2 |