Image Classification on ImageNet V2 (test)
84.3Top-1 AccuracyViT-e
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| ViT-e2022.09 | 84.3 | — | — | — | — | — | — | — | — | |
| PaLI-XResolution=756, Evaluation Protocol=fine-tuned2023.05 | 83.66 | — | — | — | — | — | — | — | — | |
| ViT-G/14Backbone=ViT-G/142021.06 | 83.33 | — | — | — | — | — | — | — | — | |
| ViT-G2022.09 | 83.3 | — | — | — | — | — | — | — | — | |
| PaLI-XResolution=224, Evaluation Protocol=fine-tuned2023.05 | 81.42 | — | — | — | — | — | — | — | — | |
| DREAM-ID2023.09 | 80.4 | — | — | — | — | — | — | — | — | |
| NSBackbone=EfficientNet-L22021.06 | 80.2 | — | — | — | — | — | — | — | — | |
| InfViT-Pure (24L)Year=Ours, Params (M)=331, Attention Type=Ours (InfSA-based)2026.02 | 79.8 | — | — | — | — | — | — | — | — | |
| AutoAugment2023.09 | 79.7 | — | — | — | — | — | — | — | — | |
| CutMix2023.09 | 79.7 | — | — | — | — | — | — | — | — | |
| AugMix2023.09 | 79.2 | — | — | — | — | — | — | — | — | |
| Generic Prompts2023.09 | 79.2 | — | — | — | — | — | — | — | — | |
| InfViT-Pure (4L)Year=Ours, Params (M)=57.7, FLOPs (G)=59.0, Attention Type=Ours (InfSA-based)2026.02 | 79.2 | — | — | — | — | — | — | — | — | |
| PaLI-17BResolution=224, Evaluation Protocol=fine-tuned2023.05 | 78.91 | — | — | — | — | — | — | — | — | |
| RandAugment2023.09 | 78.9 | — | — | — | — | — | — | — | — | |
| MEMO2023.09 | 78.6 | — | — | — | — | — | — | — | — | |
| DeepAugment2023.09 | 78.3 | — | — | — | — | — | — | — | — | |
| MAE-CTaugk=10, Backbone=ViT-L/16, Evaluation protocol=linear probing2023.04 | 77.9 | — | — | — | — | — | — | — | — | |
| Original (no aug)data augmentation=none2023.09 | 77.8 | — | — | — | — | — | — | — | — | |
| InfViT-Linear (24L)Year=Ours, Params (M)=305, Attention Type=Ours (InfSA-based)2026.02 | 77.7 | — | — | — | — | — | — | — | — | |
| InfViT-Linear (4L)Year=Ours, Params (M)=53.5, FLOPs (G)=59.0, Attention Type=Ours (InfSA-based)2026.02 | 77.4 | — | — | — | — | — | — | — | — | |
| RAVLT-LCost=Large model ~ 15.0G, Params(M)=95, FLOPS(G)=162024.11 | 76.8 | — | — | — | — | — | — | — | — | |
| RAVLT-LYear='24, Params (M)=95, FLOPs (G)=16.0, Attention Type=Linear / sub-quadratic attention2026.02 | 76.8 | — | — | — | — | — | — | — | — | |
| MLLA-BCost=Large model ~ 15.0G, Params(M)=96, FLOPS(G)=16.22024.11 | 76.7 | — | — | — | — | — | — | — | — | |
| MLLA-BYear='24, Params (M)=96, FLOPs (G)=16.2, Attention Type=Linear / sub-quadratic attention2026.02 | 76.7 | — | — | — | — | — | — | — | — | |
| Swin-B uparrow 384 (22k)Params (M)=88, MACs (B)=47.1, Input Resolution=384x3842022.04 | 76.6 | — | — | — | — | — | — | — | — | |
| MAE-CTmink=10, Backbone=ViT-L/16, Evaluation protocol=linear probing2023.04 | 76.6 | — | — | — | — | — | — | — | — | |
| MAE-CTmink=20, Backbone=ViT-L/16, Evaluation protocol=linear probing2023.04 | 76.5 | — | — | — | — | — | — | — | — | |
| MAE-CTmink=30, Backbone=ViT-L/16, Evaluation protocol=linear probing2023.04 | 76.5 | — | — | — | — | — | — | — | — | |
| MAE-CTmink=1, Backbone=ViT-L/16, Evaluation protocol=linear probing2023.04 | 76.4 | — | — | — | — | — | — | — | — | |
| RAVLT-BCost=Base model ~ 10.0G, Params(M)=48, FLOPS(G)=9.92024.11 | 76.3 | — | — | — | — | — | — | — | — | |
| RMT-LCost=Large model ~ 15.0G, Params(M)=95, FLOPS(G)=18.22024.11 | 76.3 | — | — | — | — | — | — | — | — | |
| RMT-LYear='24, Params (M)=95, FLOPs (G)=18.2, Attention Type=Softmax / standard attention2026.02 | 76.3 | — | — | — | — | — | — | — | — | |
| Mini-Swin-B uparrow 384Params (M)=47, MACs (B)=49.4, Input Resolution=384x3842022.04 | 76.1 | — | — | — | — | — | — | — | — | |
| CLIPBackbone=ViT-L/142021.06 | 75.9 | — | — | — | — | — | — | — | — | |
| RMT-BCost=Base model ~ 10.0G, Params(M)=54, FLOPS(G)=9.72024.11 | 75.6 | — | — | — | — | — | — | — | — | |
| Swin-B (22k)Params (M)=88, MACs (B)=15.4, Input Resolution=224x2242022.04 | 75.3 | — | — | — | — | — | — | — | — | |
| Mini-DeiT-B uparrow 384Params (M)=44, MACs (B)=56.9, Input Resolution=384x3842022.04 | 75.2 | — | — | — | — | — | — | — | — | |
| RAVLT-SCost=Small model (~4.5G), Params(M)=26, FLOPS(G)=4.62024.11 | 74.9 | — | — | — | — | — | — | — | — | |
| MLLA-SCost=Base model ~ 10.0G, Params(M)=43, FLOPS(G)=7.32024.11 | 74.9 | — | — | — | — | — | — | — | — | |
| DeiT-B uparrow 384Params (M)=88, MACs (B)=55.7, Input Resolution=384x3842022.04 | 74.8 | — | — | — | — | — | — | — | — | |
| Mini-Swin-BParams (M)=46, MACs (B)=15.7, Input Resolution=224x2242022.04 | 74.4 | — | — | — | — | — | — | — | — | |
| MOAT-2Cost=Large model ~ 15.0G, Params(M)=73, FLOPS(G)=17.22024.11 | 74.3 | — | — | — | — | — | — | — | — | |
| CaRotBackbone=OpenAI ViT-B/162026.05 | 74.3 | — | — | — | — | — | — | — | — | |
| MOAT-1Cost=Base model ~ 10.0G, Params(M)=42, FLOPS(G)=9.12024.11 | 74.2 | — | — | — | — | — | — | — | — | |
| RMT-SCost=Small model (~4.5G), Params(M)=27, FLOPS(G)=4.52024.11 | 74.1 | — | — | — | — | — | — | — | — | |
| ViT-B/16Year='21, Params (M)=86, FLOPs (G)=17.6, Attention Type=Softmax / standard attention2026.02 | 74.1 | — | — | — | — | — | — | — | — | |
| BiFormer-BCost=Base model ~ 10.0G, Params(M)=57, FLOPS(G)=9.82024.11 | 74 | — | — | — | — | — | — | — | — | |
| SAE-FTBackbone=OpenAI ViT-B/162026.05 | 73.9 | — | — | — | — | — | — | — | — | |
| Mini-Swin-SParams (M)=26, MACs (B)=8.9, Input Resolution=224x2242022.04 | 73.8 | — | — | — | — | — | — | — | — | |
| StarFTBackbone=OpenAI ViT-B/162026.05 | 73.8 | — | — | — | — | — | — | — | — | |
| EfficientNet-B5Params (M)=30, MACs (B)=9.9, Input Resolution=456x4562022.04 | 73.6 | — | — | — | — | — | — | — | — | |
| BiFormer-SCost=Small model (~4.5G), Params(M)=26, FLOPS(G)=4.52024.11 | 73.6 | — | — | — | — | — | — | — | — | |
| SMT-SCost=Small model (~4.5G), Params(M)=21, FLOPS(G)=4.72024.11 | 73.3 | — | — | — | — | — | — | — | — | |
| MLLA-TCost=Small model (~4.5G), Params(M)=25, FLOPS(G)=4.22024.11 | 73.3 | — | — | — | — | — | — | — | — | |
| XCIT-S24Cost=Base model ~ 10.0G, Params(M)=48, FLOPS(G)=9.12024.11 | 73.3 | — | — | — | — | — | — | — | — | |
| Swin-B uparrow 384Params (M)=88, MACs (B)=47.1, Input Resolution=384x3842022.04 | 73.2 | — | — | — | — | — | — | — | — | |
| Mini-DeiT-BParams (M)=44, MACs (B)=17.7, Input Resolution=224x2242022.04 | 73 | — | — | — | — | — | — | — | — | |
| MOAT-0Cost=Small model (~4.5G), Params(M)=28, FLOPS(G)=5.72024.11 | 72.8 | — | — | — | — | — | — | — | — | |
| CAR-FTBackbone=OpenAI ViT-B/162026.05 | 72.8 | — | — | — | — | — | — | — | — | |
| WiSE-FTBackbone=OpenAI ViT-B/162026.05 | 72.8 | — | — | — | — | — | — | — | — | |
| RAVLT-TCost=Tiny model (~2.5G), Params(M)=15, FLOPS(G)=2.42024.11 | 72.7 | — | — | — | — | — | — | — | — | |
| FLYPBackbone=OpenAI ViT-B/162026.05 | 72.7 | — | — | — | — | — | — | — | — | |
| Swin-BParams (M)=88, MACs (B)=15.4, Input Resolution=224x2242022.04 | 72.5 | — | — | — | — | — | — | — | — | |
| RegNetY-16GFParams (M)=84, MACs (B)=15.9, Input Resolution=224x2242022.04 | 72.4 | — | — | — | — | — | — | — | — | |
| DeiT-B uparrow 384Params (M)=87, MACs (B)=55.6, Input Resolution=384x3842022.04 | 72.4 | — | — | — | — | — | — | — | — | |
| RMT-TCost=Tiny model (~2.5G), Params(M)=14, FLOPS(G)=2.52024.11 | 72.1 | — | — | — | — | — | — | — | — | |
| ECA-Resnet269-DTraining procedure=A12021.10 | 71.9 | — | — | — | — | — | — | — | — | |
| Swin-SParams (M)=50, MACs (B)=8.7, Input Resolution=224x2242022.04 | 71.9 | — | — | — | — | — | — | — | — | |
| RegNetY-32GFTraining procedure=A12021.10 | 71.7 | — | — | — | — | — | — | — | — | |
| EfficientNetV2-rw-MTraining procedure=A12021.10 | 71.7 | — | — | — | — | — | — | — | — | |
| MAEk=-, Backbone=ViT-L/16, Evaluation protocol=linear probing2023.04 | 71.7 | — | — | — | — | — | — | — | — | |
| FTBackbone=OpenAI ViT-B/162026.05 | 71.7 | — | — | — | — | — | — | — | — | |
| DeiT-BParams (M)=86, MACs (B)=17.6, Input Resolution=224x2242022.04 | 71.5 | — | — | — | — | — | — | — | — | |
| DeiT-BCost=Large model ~ 15.0G, Params(M)=86, FLOPS(G)=17.52024.11 | 71.5 | — | — | — | — | — | — | — | — | |
| DeiT-BYear='21, Params (M)=86, FLOPs (G)=17.5, Attention Type=Softmax / standard attention2026.02 | 71.5 | — | — | — | — | — | — | — | — | |
| RegNetY-16GFTraining procedure=A12021.10 | 71.2 | — | — | — | — | — | — | — | — | |
| SENet-154Training procedure=A12021.10 | 71.2 | — | — | — | — | — | — | — | — | |
| RegNetY-8GFTraining procedure=A12021.10 | 71.1 | — | — | — | — | — | — | — | — | |
| SMT-TCost=Tiny model (~2.5G), Params(M)=12, FLOPS(G)=2.42024.11 | 71 | — | — | — | — | — | — | — | — | |
| EfficientNet-B4Training procedure=A12021.10 | 70.8 | — | — | — | — | — | — | — | — | |
| RegNetY-4GFTraining procedure=A12021.10 | 70.7 | — | — | — | — | — | — | — | — | |
| BiFormer-TCost=Tiny model (~2.5G), Params(M)=13, FLOPS(G)=2.22024.11 | 70.7 | — | — | — | — | — | — | — | — | |
| ResNet-152Training procedure=A12021.10 | 70.6 | — | — | — | — | — | — | — | — | |
| Mini-Swin-TParams (M)=12, MACs (B)=4.6, Input Resolution=224x2242022.04 | 70.5 | — | — | — | — | — | — | — | — | |
| EfficientNet-B3Training procedure=A12021.10 | 70.4 | — | — | — | — | — | — | — | — | |
| ResNet-101Training procedure=A12021.10 | 70.3 | — | — | — | — | — | — | — | — | |
| ALIGNBackbone=EfficientNet-L22021.06 | 70.1 | — | — | — | — | — | — | — | — | |
| ALIGNTraining Dataset=ALIGN-1800M, Evaluation Protocol=Zero-shot, Resolution=640x640, Implementation Source=Official2022.10 | 70.1 | — | — | — | — | — | — | — | — | |
| tiny-MOAT-2Cost=Tiny model (~2.5G), Params(M)=10, FLOPS(G)=2.32024.11 | 70.1 | — | — | — | — | — | — | — | — | |
| ECA-ResNet50-TTraining procedure=A12021.10 | 69.9 | — | — | — | — | — | — | — | — | |
| Swin-TParams (M)=28, MACs (B)=4.5, Input Resolution=224x2242022.04 | 69.6 | — | — | — | — | — | — | — | — | |
| Mini-DeiT-SParams (M)=11, MACs (B)=4.7, Input Resolution=224x2242022.04 | 69.5 | — | — | — | — | — | — | — | — | |
| ViT-STraining procedure=A12021.10 | 69.4 | — | — | — | — | — | — | — | — | |
| ViT-BTraining procedure=A12021.10 | 69.4 | — | — | — | — | — | — | — | — | |
| EfficientNet-B2Training procedure=A12021.10 | 69.3 | — | — | — | — | — | — | — | — | |
| EfficientNetV2-rw-STraining procedure=A12021.10 | 69.2 | — | — | — | — | — | — | — | — | |
| TuneCLIPBase Model=SigLIP ViT-B/162026.01 | 69.02 | — | — | — | — | — | — | — | — | |
| BaselineBase Model=SigLIP ViT-B/162026.01 | 68.93 | — | — | — | — | — | — | — | — | |
| ResNet-50-DTraining procedure=A12021.10 | 68.9 | — | — | — | — | — | — | — | — |