Video Recognition on Kinetics-400 (test)
90.25ASRI2V
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| I2VSource Image Model=Resnet-101, Target Video Model=TPN-502021.12 | 90.25 | — | — | — | |
| ENS-I2VSource Image Model=Ensemble, Target Video Model=TPN-502021.12 | 88 | — | — | — | |
| I2VSource Image Model=Resnet-101, Target Video Model=TPN-1012021.12 | 87.25 | — | — | — | |
| ENS-I2VSource Image Model=Ensemble, Target Video Model=TPN-1012021.12 | 85.75 | — | — | — | |
| ENS-I2VSource Image Model=Ensemble, Target Video Model=SlowFast-1012021.12 | 79.75 | — | — | — | |
| I2VSource Image Model=Resnet-101, Target Video Model=SlowFast-502021.12 | 77 | — | — | — | |
| ENS-I2VSource Image Model=Ensemble, Target Video Model=SlowFast-502021.12 | 76.5 | — | — | — | |
| I2VSource Image Model=Resnet-101, Target Video Model=SlowFast-1012021.12 | 74.75 | — | — | — | |
| ENS-I2VSource Image Model=Ensemble, Target Video Model=NL-502021.12 | 72.25 | — | — | — | |
| I2VSource Image Model=Vgg-16, Target Video Model=TPN-502021.12 | 70.5 | — | — | — | |
| I2VSource Image Model=Alexnet, Target Video Model=TPN-502021.12 | 69.5 | — | — | — | |
| ENS-I2VSource Image Model=Ensemble, Target Video Model=NL-1012021.12 | 65 | — | — | — | |
| I2VSource Image Model=Resnet-101, Target Video Model=NL-502021.12 | 64.5 | — | — | — | |
| I2VSource Image Model=Squeezenet, Target Video Model=SlowFast-1012021.12 | 62.5 | — | — | — | |
| I2VSource Image Model=Alexnet, Target Video Model=SlowFast-1012021.12 | 61.5 | — | — | — | |
| I2VSource Image Model=Squeezenet, Target Video Model=SlowFast-502021.12 | 60.25 | — | — | — | |
| I2VSource Image Model=Alexnet, Target Video Model=TPN-1012021.12 | 59.75 | — | — | — | |
| I2VSource Image Model=Alexnet, Target Video Model=SlowFast-502021.12 | 59.5 | — | — | — | |
| I2VSource Image Model=Vgg-16, Target Video Model=SlowFast-502021.12 | 59 | — | — | — | |
| I2VSource Image Model=Vgg-16, Target Video Model=TPN-1012021.12 | 59 | — | — | — | |
| I2VSource Image Model=Squeezenet, Target Video Model=TPN-502021.12 | 58.5 | — | — | — | |
| I2VSource Image Model=Vgg-16, Target Video Model=SlowFast-1012021.12 | 57.75 | — | — | — | |
| I2VSource Image Model=Resnet-101, Target Video Model=NL-1012021.12 | 56.25 | — | — | — | |
| I2VSource Image Model=Squeezenet, Target Video Model=TPN-1012021.12 | 55.5 | — | — | — | |
| I2VSource Image Model=Alexnet, Target Video Model=NL-502021.12 | 54.75 | — | — | — | |
| DRSource Image Model=Resnet-101, Target Video Model=SlowFast-502021.12 | 52.25 | — | — | — | |
| I2VSource Image Model=Squeezenet, Target Video Model=NL-502021.12 | 51 | — | — | — | |
| DRSource Image Model=Resnet-101, Target Video Model=SlowFast-1012021.12 | 49 | — | — | — | |
| I2VSource Image Model=Vgg-16, Target Video Model=NL-502021.12 | 46.25 | — | — | — | |
| I2VSource Image Model=Alexnet, Target Video Model=NL-1012021.12 | 44 | — | — | — | |
| DRSource Image Model=Alexnet, Target Video Model=SlowFast-1012021.12 | 43 | — | — | — | |
| DRSource Image Model=Resnet-101, Target Video Model=TPN-502021.12 | 42.75 | — | — | — | |
| DRSource Image Model=Alexnet, Target Video Model=SlowFast-502021.12 | 41.75 | — | — | — | |
| DRSource Image Model=Resnet-101, Target Video Model=TPN-1012021.12 | 41.5 | — | — | — | |
| DRSource Image Model=Alexnet, Target Video Model=TPN-502021.12 | 39 | — | — | — | |
| I2VSource Image Model=Vgg-16, Target Video Model=NL-1012021.12 | 39 | — | — | — | |
| I2VSource Image Model=Squeezenet, Target Video Model=NL-1012021.12 | 37.75 | — | — | — | |
| DRSource Image Model=Resnet-101, Target Video Model=NL-502021.12 | 37.25 | — | — | — | |
| DRSource Image Model=Squeezenet, Target Video Model=SlowFast-1012021.12 | 37 | — | — | — | |
| DRSource Image Model=Vgg-16, Target Video Model=SlowFast-1012021.12 | 36.75 | — | — | — | |
| DRSource Image Model=Squeezenet, Target Video Model=SlowFast-502021.12 | 36.5 | — | — | — | |
| DRSource Image Model=Vgg-16, Target Video Model=SlowFast-502021.12 | 35.75 | — | — | — | |
| DRSource Image Model=Alexnet, Target Video Model=NL-502021.12 | 31.5 | — | — | — | |
| DRSource Image Model=Alexnet, Target Video Model=TPN-1012021.12 | 31 | — | — | — | |
| DRSource Image Model=Squeezenet, Target Video Model=TPN-502021.12 | 29.5 | — | — | — | |
| DRSource Image Model=Vgg-16, Target Video Model=TPN-502021.12 | 29 | — | — | — | |
| DRSource Image Model=Resnet-101, Target Video Model=NL-1012021.12 | 25.5 | — | — | — | |
| DRSource Image Model=Squeezenet, Target Video Model=NL-502021.12 | 25 | — | — | — | |
| DRSource Image Model=Squeezenet, Target Video Model=TPN-1012021.12 | 24.25 | — | — | — | |
| DRSource Image Model=Vgg-16, Target Video Model=TPN-1012021.12 | 23.75 | — | — | — | |
| DRSource Image Model=Vgg-16, Target Video Model=NL-502021.12 | 23 | — | — | — | |
| DRSource Image Model=Alexnet, Target Video Model=NL-1012021.12 | 22 | — | — | — | |
| DRSource Image Model=Squeezenet, Target Video Model=NL-1012021.12 | 17 | — | — | — | |
| DRSource Image Model=Vgg-16, Target Video Model=NL-1012021.12 | 16.75 | — | — | — | |
| ActionCLIP-B/16Evaluation Protocol=Adapting, Inputs (frames×crops×clips)=32×10×3, GFLOPs=563, Backbone=ViT-B/162026.05 | — | 83.8 | 96.2 | — | |
| AIM-B/16Evaluation Protocol=Adapting, Inputs (frames×crops×clips)=8×1×3, GFLOPs=202, Backbone=ViT-B/162026.05 | — | 83.9 | 96.3 | — | |
| AIM-L/14Evaluation Protocol=Adapting, Inputs (frames×crops×clips)=8×1×3, GFLOPs=934, Backbone=ViT-L/142026.05 | — | 86.8 | 97.2 | — | |
| DUALPATH-L/14Evaluation Protocol=Adapting, Inputs (frames×crops×clips)=32×1×3, Backbone=ViT-L/142026.05 | — | 87.7 | 97.8 | — | |
| ETL-ViCLIP-B/16Evaluation Protocol=Adapting, Inputs (frames×crops×clips)=8×4×3, Backbone=ViT-B/162026.05 | — | 82.2 | 96.2 | — | |
| EVL ViT-B/16Pretraining=CLIP, #Frames=8 x 3, GFLOPS=4442022.08 | — | 82.9 | — | — | |
| EVL ViT-B/16Pretraining=CLIP, #Frames=16 x 3, GFLOPS=8882022.08 | — | 83.6 | — | — | |
| EVL ViT-B/16Pretraining=CLIP, #Frames=32 x 3, GFLOPS=17772022.08 | — | 84.2 | — | — | |
| EVL ViT-L/14Pretraining=CLIP, #Frames=8 x 3, GFLOPS=20222022.08 | — | 86.3 | — | — | |
| EVL ViT-L/14Pretraining=CLIP, #Frames=16 x 3, GFLOPS=40442022.08 | — | 87 | — | — | |
| EVL ViT-L/14Pretraining=CLIP, #Frames=32 x 3, GFLOPS=80882022.08 | — | 87.3 | — | — | |
| EVL ViT-L/14Pretraining=CLIP, Resolution=336px, #Frames=32 x 3, GFLOPS=181962022.08 | — | 87.7 | — | — | |
| FocusVideo-B/16Evaluation Protocol=Adapting, Inputs (frames×crops×clips)=8×4×3, GFLOPs=204, Backbone=ViT-B/162026.05 | — | 84.1 | 96.5 | — | |
| FocusVideo-B/16Evaluation Protocol=Adapting, Inputs (frames×crops×clips)=32×4×3, GFLOPs=816, Backbone=ViT-B/162026.05 | — | 84.7 | 96.8 | — | |
| FocusVideo-L/14Evaluation Protocol=Adapting, Inputs (frames×crops×clips)=8×4×3, GFLOPs=914, Backbone=ViT-L/142026.05 | — | 87.2 | 97.7 | — | |
| FocusVideo-L/14Evaluation Protocol=Adapting, Inputs (frames×crops×clips)=32×4×3, GFLOPs=3656, Backbone=ViT-L/142026.05 | — | 88 | 97.9 | — | |
| FoPEResolution=160, YaRN=false, Backbone=TimeSformer, Training Resolution=224 x 2242026.06 | — | — | — | 69.93 | |
| FoPE + YaRNResolution=1024, YaRN=true, Backbone=TimeSformer, Training Resolution=224 x 2242026.06 | — | — | — | 55.87 | |
| irCSN-152Pretraining=IG-65M, #Frames=32 x 30, GFLOPS=29012022.08 | — | 82.6 | — | — | |
| Learnable PEResolution=160, YaRN=false, Backbone=TimeSformer, Training Resolution=224 x 2242026.06 | — | — | — | 70.05 | |
| M2-CLIP-B/16Evaluation Protocol=Adapting, Inputs (frames×crops×clips)=8×4×3, GFLOPs=214, Backbone=ViT-B/162026.05 | — | 83.4 | 96.3 | — | |
| M2-CLIP-B/16Evaluation Protocol=Adapting, Inputs (frames×crops×clips)=32×4×3, GFLOPs=842, Backbone=ViT-B/162026.05 | — | 84.1 | 96.8 | — | |
| M2-CLIP-L/14Evaluation Protocol=Adapting, Inputs (frames×crops×clips)=32×4×3, Backbone=ViT-L/142026.05 | — | 87 | 97.6 | — | |
| MoTE-B/16Evaluation Protocol=Tuning, Inputs (frames×crops×clips)=8×4×3, GFLOPs=141, Backbone=ViT-B/162026.05 | — | 83 | 96.3 | — | |
| MoTE-L/14Evaluation Protocol=Tuning, Inputs (frames×crops×clips)=8×4×3, GFLOPs=649, Backbone=ViT-L/142026.05 | — | 86.8 | 97.5 | — | |
| MoTE-L/14Evaluation Protocol=Tuning, Inputs (frames×crops×clips)=16×4×3, GFLOPs=1299, Backbone=ViT-L/142026.05 | — | 87.2 | 97.7 | — | |
| MTV-LPretraining=JFT, #Frames=32 x 12, GFLOPS=180502022.08 | — | 84.3 | — | — | |
| MViT-LPretraining=MaskFeat, K600, #Frames=16 x 10, GFLOPS=37702022.08 | — | 85.1 | — | — | |
| MViT-SPretraining=ImageNet-21k, #Frames=16 x 10, GFLOPS=7102022.08 | — | 82.6 | — | — | |
| nD-RoPEResolution=160, YaRN=false, Backbone=TimeSformer, Training Resolution=224 x 2242026.06 | — | — | — | 70.63 | |
| nD-RoPEResolution=224, YaRN=false, Backbone=TimeSformer, Training Resolution=224 x 2242026.06 | — | — | — | 75.85 | |
| nD-RoPEResolution=512, YaRN=false, Backbone=TimeSformer, Training Resolution=224 x 2242026.06 | — | — | — | 67.48 | |
| nD-RoPEResolution=1024, YaRN=false, Backbone=TimeSformer, Training Resolution=224 x 2242026.06 | — | — | — | 29.21 | |
| nD-RoPE + YaRNResolution=224, YaRN=true, Backbone=TimeSformer, Training Resolution=224 x 2242026.06 | — | — | — | 75.85 | |
| nD-RoPE + YaRNResolution=512, YaRN=true, Backbone=TimeSformer, Training Resolution=224 x 2242026.06 | — | — | — | 72.91 | |
| nD-RoPE + YaRNResolution=1024, YaRN=true, Backbone=TimeSformer, Training Resolution=224 x 2242026.06 | — | — | — | 59.23 | |
| Omnivore-BPretraining=IN1k + SUN, #Frames=32 x 12, GFLOPS=33842022.08 | — | 83.3 | — | — | |
| OST-B/16Evaluation Protocol=Tuning, Inputs (frames×crops×clips)=16×1×1, Backbone=ViT-B/162026.05 | — | 83.2 | — | — | |
| RoPE-AxialResolution=160, YaRN=false, Backbone=TimeSformer, Training Resolution=224 x 2242026.06 | — | — | — | 66.25 | |
| RoPE-Axial + APEResolution=160, YaRN=false, Backbone=TimeSformer, Training Resolution=224 x 2242026.06 | — | — | — | 66.24 | |
| RoPE-MixedResolution=160, YaRN=false, Backbone=TimeSformer, Training Resolution=224 x 2242026.06 | — | — | — | 66.76 | |
| RoPE-Mixed + APEResolution=160, YaRN=false, Backbone=TimeSformer, Training Resolution=224 x 2242026.06 | — | — | — | 66.47 | |
| ST-Adapter-B/16Evaluation Protocol=Adapting, Inputs (frames×crops×clips)=8×1×3, GFLOPs=148, Backbone=ViT-B/162026.05 | — | 82 | 95.7 | — | |
| ST-Adapter-B/16Evaluation Protocol=Adapting, Inputs (frames×crops×clips)=32×1×3, GFLOPs=607, Backbone=ViT-B/162026.05 | — | 82.7 | 96.2 | — | |
| ST-Adapter-L/14Evaluation Protocol=Adapting, Inputs (frames×crops×clips)=8×1×3, GFLOPs=687, Backbone=ViT-L/142026.05 | — | 86.7 | 97.5 | — | |
| ST-Adapter-L/14Evaluation Protocol=Adapting, Inputs (frames×crops×clips)=32×1×3, GFLOPs=2749, Backbone=ViT-L/142026.05 | — | 87.2 | 97.6 | — |