Image Classification on ImageNet V2
86.9Top-1 AccViT-B/16
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| ViT-B/16GFLOPs=16.92025.10 | 86.9 | — | — | — | — | — | — | — | |
| ViT-L/16GFLOPs=59.72025.10 | 86 | — | — | — | — | — | — | — | |
| Best model on each test set (oracle)Backbone=BASIC-L, Evaluation Protocol=fine-tuned, Model Selection Strategy=oracle (best on test set)2022.03 | 84.84 | — | — | — | — | — | — | — | |
| ResNet101GFLOPs=7.82025.10 | 84.8 | — | — | — | — | — | — | — | |
| Greedy ensembleBackbone=BASIC-L, Evaluation Protocol=fine-tuned, Model Selection Strategy=greedy ensemble2022.03 | 84.65 | — | — | — | — | — | — | — | |
| 22B emaModel backbone=22B, Fine-tuned resolution=560px, Polyak averaging (EMA)=true, Training setup=Fine-tuned on ImageNet2023.02 | 84.65 | — | — | — | — | — | — | — | |
| Greedy soupBackbone=BASIC-L, Evaluation Protocol=fine-tuned, Model Selection Strategy=greedy soup2022.03 | 84.63 | — | — | — | — | — | — | — | |
| Best model on held out val setBackbone=BASIC-L, Evaluation Protocol=fine-tuned, Model Selection Strategy=best on val set2022.03 | 84.42 | — | — | — | — | — | — | — | |
| e/14 emaModel backbone=e/14, Fine-tuned resolution=560px, Polyak averaging (EMA)=true, Training setup=Fine-tuned on ImageNet2023.02 | 84.38 | — | — | — | — | — | — | — | |
| 22BModel backbone=22B, Fine-tuned resolution=560px, Polyak averaging (EMA)=false, Training setup=Fine-tuned on ImageNet2023.02 | 84.38 | — | — | — | — | — | — | — | |
| ViT-e/14Evaluation Protocol=High-res fine-tuning2023.02 | 84.3 | — | — | — | — | — | — | — | |
| 22BModel backbone=22B, Fine-tuned resolution=384px, Polyak averaging (EMA)=false, Training setup=Fine-tuned on ImageNet2023.02 | 84.28 | — | — | — | — | — | — | — | |
| ViT-G/14 greedy soupBackbone=ViT/G-14, Model Selection Strategy=greedy soup2022.03 | 84.22 | — | — | — | — | — | — | — | |
| Model Soups ViT-GBackbone=ViT-G2025.07 | 84.22 | — | — | — | — | — | — | — | |
| SwinV2Pre-training Dataset=IN-ext-70M, Architecture=SwinV2-G, Resolution=6402023.03 | 84 | — | — | — | — | — | — | — | |
| MAWSPre-training Dataset=IG-3B, Architecture=ViT-6.5B, Resolution=5182023.03 | 84 | — | — | — | — | — | — | — | |
| e/14 emaModel backbone=e/14, Fine-tuned resolution=384px, Polyak averaging (EMA)=true, Training setup=Fine-tuned on ImageNet2023.02 | 83.95 | — | — | — | — | — | — | — | |
| APMP (Pre-trained on clean ImageNet)=false, Backbone=VIT-H/142024.10 | 83.9 | — | — | — | — | — | — | — | |
| ViT-L/32GFLOPs=15.32025.10 | 83.9 | — | — | — | — | — | — | — | |
| PaLI-XResolution=756, Setting=Fine-tuning, Steps=2.2x more steps2023.05 | 83.66 | — | — | — | — | — | — | — | |
| g/14 emaModel backbone=g/14, Fine-tuned resolution=518px, Polyak averaging (EMA)=true, Training setup=Fine-tuned on ImageNet2023.02 | 83.61 | — | — | — | — | — | — | — | |
| G/14 emaModel backbone=G/14, Fine-tuned resolution=518px, Polyak averaging (EMA)=true, Training setup=Fine-tuned on ImageNet2023.02 | 83.53 | — | — | — | — | — | — | — | |
| PaLI-XResolution=756, Setting=Fine-tuning2023.05 | 83.48 | — | — | — | — | — | — | — | |
| ResNet50GFLOPs=4.12025.10 | 83.4 | — | — | — | — | — | — | — | |
| LionModel=ViT-g/14, Input Resolution=518x518, #Params=1.04B, Pre-training Dataset=JFT-3B2023.02 | 83.39 | — | — | — | — | — | — | — | |
| ViT/G-14Backbone=ViT/G-142022.03 | 83.33 | — | — | — | — | — | — | — | |
| ViT-G/14Evaluation Protocol=High-res fine-tuning2023.02 | 83.33 | — | — | — | — | — | — | — | |
| AdafactorModel=ViT-G/14, Input Resolution=518x518, #Params=1.88B, Pre-training Dataset=JFT-3B2023.02 | 83.33 | — | — | — | — | — | — | — | |
| ViT G/14Pre-training=JFT 3B, Fine-tuning protocol=Finetuned on ImageNet-1k2022.01 | 83.3 | — | — | — | — | — | — | — | |
| Scale-ViTPre-training Dataset=JFT-3B, Architecture=ViT-G, Resolution=5182023.03 | 83.3 | — | — | — | — | — | — | — | |
| ViT-22BParameters=21.7B, Evaluation Protocol=Linear Probing, Training Data=JFT-3B2023.12 | 83.2 | — | — | — | — | — | — | — | |
| ViT-22BEvaluation Protocol=Linear Probing (frozen), Resolution=224px2023.02 | 83.15 | — | — | — | — | — | — | — | |
| AdafactorModel=ViT-g/14, Input Resolution=518x518, #Params=1.04B, Pre-training Dataset=JFT-3B2023.02 | 83.1 | — | — | — | — | — | — | — | |
| MAWSPre-training Dataset=IG-3B, Architecture=ViT-2B, Resolution=5182023.03 | 83 | — | — | — | — | — | — | — | |
| ViT-e/14Evaluation Protocol=Linear Probing (frozen), Resolution=224px2023.02 | 82.51 | — | — | — | — | — | — | — | |
| MAWSPre-training Dataset=IG-3B, Architecture=ViT-H, Resolution=5182023.03 | 82.3 | — | — | — | — | — | — | — | |
| LionModel=ViT-H/14, Input Resolution=518x518, #Params=633.47M, Pre-training Dataset=JFT-300M2023.02 | 82.24 | — | — | — | — | — | — | — | |
| PaLI-XResolution=224, Setting=Fine-tuning2023.05 | 81.42 | — | — | — | — | — | — | — | |
| ViT-H/14Optimizer=Lion, Parameters=632.72M, Epochs / Steps=14 / 1,035,583, Fine-tuning resolution=392x392, Pre-training=JFT-300M, Polyak averaging=false2023.02 | 81.41 | — | — | — | — | — | — | — | |
| ViT-G/14Evaluation Protocol=Linear Probing (frozen), Resolution=224px2023.02 | 81.32 | — | — | — | — | — | — | — | |
| ResNet34GFLOPs=3.72025.10 | 81.3 | — | — | — | — | — | — | — | |
| LionBackbone=BASIC-L, Evaluation Protocol=Zero-shot2023.02 | 81.2 | — | — | — | — | — | — | — | |
| LionModel=ViT-L/16, Input Resolution=512x512, #Params=305.18M, Pre-training Dataset=JFT-300M2023.02 | 81.13 | — | — | — | — | — | — | — | |
| AdamWModel=ViT-H/14, Input Resolution=518x518, #Params=633.47M, Pre-training Dataset=JFT-300M2023.02 | 81.12 | — | — | — | — | — | — | — | |
| ViT H/14Pre-training=IG 3.6B, Fine-tuning protocol=Finetuned on ImageNet-1k2022.01 | 81.1 | — | — | — | — | — | — | — | |
| ViT-g/14Evaluation Protocol=Linear Probing (frozen), Resolution=224px2023.02 | 81.1 | — | — | — | — | — | — | — | |
| SWAGPre-training Dataset=IG-3.6B, Architecture=ViT-H, Resolution=5182023.03 | 81.1 | — | — | — | — | — | — | — | |
| FixNoisy-L2Evaluation Protocol=High-res fine-tuning2023.02 | 80.8 | — | — | — | — | — | — | — | |
| L/16Model backbone=L/16, Fine-tuned resolution=384px, Polyak averaging (EMA)=false, Training setup=Fine-tuned on ImageNet2023.02 | 80.74 | — | — | — | — | — | — | — | |
| OpenCLIP-VIT-H/14P (Pre-trained on clean ImageNet)=false, Backbone=VIT-H/142024.10 | 80.7 | — | — | — | — | — | — | — | |
| CoCaEvaluation Protocol=Zero-shot2022.05 | 80.7 | — | — | — | — | — | — | — | |
| CoCaBackbone=CoCa, Evaluation Protocol=zero-shot2022.03 | 80.7 | — | — | — | — | — | — | — | |
| CoCaZero-shot=true2022.09 | 80.7 | — | — | — | — | — | — | — | |
| BASICEvaluation Protocol=Zero-shot2022.05 | 80.6 | — | — | — | — | — | — | — | |
| BASIC-LBackbone=BASIC-L, Evaluation Protocol=zero-shot2022.03 | 80.6 | — | — | — | — | — | — | — | |
| BASICZero-shot=true2021.11 | 80.6 | — | — | — | — | — | — | — | |
| AdafactorBackbone=BASIC-L, Evaluation Protocol=Zero-shot2023.02 | 80.6 | — | — | — | — | — | — | — | |
| BASICZero-shot=true2022.09 | 80.6 | — | — | — | — | — | — | — | |
| LiT ViT-eZero-shot=true2022.09 | 80.6 | — | — | — | — | — | — | — | |
| DC(25, 250)GFLOPs=415056.02025.10 | 80.6 | — | — | — | — | — | — | — | |
| ReLabelNetwork=ResNet-50, Supervision=ReLabel2021.01 | 80.5 | — | — | — | — | — | — | — | |
| ViT-L/16Optimizer=Lion, Parameters=304.72M, Epochs / Steps=14 / 1,035,583, Fine-tuning resolution=384x384, Pre-training=JFT-300M, Polyak averaging=false2023.02 | 80.48 | — | — | — | — | — | — | — | |
| ViT L/16Pre-training=JFT 3B, Fine-tuning protocol=Finetuned on ImageNet-1k2022.01 | 80.4 | — | — | — | — | — | — | — | |
| RegNetY 128GFPre-training=IG 3.6B, Fine-tuning protocol=Finetuned on ImageNet-1k2022.01 | 80.4 | — | — | — | — | — | — | — | |
| ViT-L/16Evaluation Protocol=High-res fine-tuning2023.02 | 80.4 | — | — | — | — | — | — | — | |
| Scale-ViTPre-training Dataset=JFT-3B, Architecture=ViT-L, Resolution=3842023.03 | 80.4 | — | — | — | — | — | — | — | |
| ViT L/16Pre-training=IG 3.6B, Fine-tuning protocol=Finetuned on ImageNet-1k2022.01 | 80.3 | — | — | — | — | — | — | — | |
| NS EfficientNet-L22021.02 | 80.2 | — | — | — | — | — | — | — | |
| EfficientNet L2Pre-training=JFT 300M+, Fine-tuning protocol=Finetuned on ImageNet-1k2022.01 | 80.2 | — | — | — | — | — | — | — | |
| ViT-H/14Optimizer=AdamW, Parameters=632.72M, Epochs / Steps=14 / 1,035,583, Fine-tuning resolution=392x392, Pre-training=JFT-300M, Polyak averaging=false2023.02 | 80.1 | — | — | — | — | — | — | — | |
| ViT-H – cosubnb params=632.1M, throughput=112 im/s, FLOPs=167.4G, peak mem=6984 MB, Pre-training=ImageNet-21k, Resolution=224, Fine-tuning epochs=502022.12 | 80 | — | — | — | — | — | — | — | |
| InternViT-6BParameters=5.9B, Evaluation Protocol=Linear Probing2023.12 | 79.9 | — | — | — | — | — | — | — | |
| LiTTraining Data Source=Private, Evaluation Protocol=Zero-shot2021.11 | 79.8 | — | — | — | — | — | — | — | |
| AdamWModel=ViT-L/16, Input Resolution=512x512, #Params=305.18M, Pre-training Dataset=JFT-300M2023.02 | 79.8 | — | — | — | — | — | — | — | |
| LiT ViT-gZero-shot=true2022.09 | 79.8 | — | — | — | — | — | — | — | |
| SigLIP2-g/16Params=1B, Zero-shot=true2026.03 | 79.8 | — | — | — | — | — | — | — | |
| CoCa-LargeEvaluation Protocol=Zero-shot2022.05 | 79.6 | — | — | — | — | — | — | — | |
| Label smoothingNetwork=ResNet-50, Supervision=Label smoothing (ε=0.1)2021.01 | 79.5 | — | — | — | — | — | — | — | |
| WiSE-FTModel=CLIP ViT-L/14@336px, Evaluation Protocol=End-to-End, Mixing Coefficient (alpha)=0.52021.09 | 79.5 | — | — | — | — | — | — | — | |
| WiSE-FTModel=CLIP ViT-L/14@336px, Evaluation Protocol=End-to-End, Mixing Coefficient (alpha)=optimal2021.09 | 79.5 | — | — | — | — | — | — | — | |
| DINOv3-LResolution=512, Evaluation Protocol=Frozen backbone + Linear head2026.05 | 79.5 | — | — | — | — | — | — | — | |
| DINOv3-LResolution=768, Evaluation Protocol=Frozen backbone + Linear head2026.05 | 79.41 | — | — | — | — | — | — | — | |
| A-VARC+GFLOPs=4649.42025.10 | 79.3 | — | — | — | — | — | — | — | |
| ViT-L/16Optimizer=Lion, Parameters=304.72M, Epochs / Steps=7 / 517,791, Fine-tuning resolution=384x384, Pre-training=JFT-300M, Polyak averaging=false2023.02 | 79.29 | — | — | — | — | — | — | — | |
| DINOv3-LResolution=384, Evaluation Protocol=Frozen backbone + Linear head2026.05 | 79.28 | — | — | — | — | — | — | — | |
| CaRotBackbone=CLIP ViT-L/142026.05 | 79.28 | — | — | — | — | — | 0.0634 | — | |
| ViT-H – DeiT-IIInb params=632.1M, throughput=112 im/s, FLOPs=167.4G, peak mem=6984 MB, Pre-training=ImageNet-21k, Resolution=2242022.12 | 79.2 | — | — | — | — | — | — | — | |
| Label cleaningNetwork=ResNet-50, Supervision=Label cleaning2021.01 | 79.1 | — | — | — | — | — | — | — | |
| ViT-L – cosubnb params=304.4M, throughput=277 im/s, FLOPs=61.6G, peak mem=3789 MB, Pre-training=ImageNet-21k, Resolution=224, Fine-tuning epochs=502022.12 | 79.1 | — | — | — | — | — | — | — | |
| ResNet18GFLOPs=1.82025.10 | 79.1 | — | — | — | — | — | — | — | |
| OriginalNetwork=ResNet-50, Supervision=Original2021.01 | 79 | — | — | — | — | — | — | — | |
| ViT-L/16Optimizer=AdamW, Parameters=304.72M, Epochs / Steps=14 / 1,035,583, Fine-tuning resolution=384x384, Pre-training=JFT-300M, Polyak averaging=false2023.02 | 78.91 | — | — | — | — | — | — | — | |
| PaLIParameters=17B, Resolution=224, Setting=Fine-tuning2023.05 | 78.91 | — | — | — | — | — | — | — | |
| DINOv3-LResolution=256, Evaluation Protocol=Frozen backbone + Linear head2026.05 | 78.84 | — | — | — | — | — | — | — | |
| MAE-H + DAT2022.09 | 78.82 | — | — | — | — | — | — | — | |
| LiTEvaluation Protocol=Zero-shot2022.05 | 78.7 | — | — | — | — | — | — | — | |
| SigLIP2-so400m/14Params=0.4B, Zero-shot=true2026.03 | 78.7 | — | — | — | — | — | — | — | |
| ViT-L – DeiT-IIInb params=304.4M, throughput=277 im/s, FLOPs=61.6G, peak mem=3789 MB, Pre-training=ImageNet-21k, Resolution=2242022.12 | 78.6 | — | — | — | — | — | — | — | |
| ViT-L/16Evaluation Protocol=Linear Probing (frozen), Resolution=224px2023.02 | 78.57 | — | — | — | — | — | — | — | |
| TRACERBackbone=CLIP ViT-L/142026.05 | 78.54 | — | — | — | — | — | 0.0581 | — |