Image Classification on ImageNet-R
96.8Top-1 AccLion
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| LionBackbone=BASIC-L, Evaluation Protocol=Zero-shot2023.02 | 96.8 | — | — | — | — | — | |
| CoCaEvaluation Protocol=Zero-shot2022.05 | 96.5 | — | — | — | — | — | |
| CoCaBackbone=CoCa, Evaluation Protocol=zero-shot2022.03 | 96.5 | — | — | — | — | — | |
| CoCaZero-shot=true2022.09 | 96.5 | — | — | — | — | — | |
| Greedy soupBackbone=BASIC-L, Evaluation Protocol=fine-tuned, Model Selection Strategy=greedy soup2022.03 | 96.1 | — | — | — | — | — | |
| LiT ViT-eZero-shot=true2022.09 | 96.1 | — | — | — | — | — | |
| Best model on each test set (oracle)Backbone=BASIC-L, Evaluation Protocol=fine-tuned, Model Selection Strategy=oracle (best on test set)2022.03 | 95.89 | — | — | — | — | — | |
| Greedy ensembleBackbone=BASIC-L, Evaluation Protocol=fine-tuned, Model Selection Strategy=greedy ensemble2022.03 | 95.79 | — | — | — | — | — | |
| BASICEvaluation Protocol=Zero-shot2022.05 | 95.7 | — | — | — | — | — | |
| BASIC-LBackbone=BASIC-L, Evaluation Protocol=zero-shot2022.03 | 95.7 | — | — | — | — | — | |
| BASICZero-shot=true2021.11 | 95.7 | — | — | — | — | — | |
| AdafactorBackbone=BASIC-L, Evaluation Protocol=Zero-shot2023.02 | 95.7 | — | — | — | — | — | |
| BASICZero-shot=true2022.09 | 95.7 | — | — | — | — | — | |
| CoCa-LargeEvaluation Protocol=Zero-shot2022.05 | 95.6 | — | — | — | — | — | |
| Best model on held out val setBackbone=BASIC-L, Evaluation Protocol=fine-tuned, Model Selection Strategy=best on val set2022.03 | 95.5 | — | — | — | — | — | |
| ViT-G/14 greedy soupBackbone=ViT/G-14, Model Selection Strategy=greedy soup2022.03 | 95.46 | — | — | — | — | — | |
| Model Soups ViT-GBackbone=ViT-G2025.07 | 95.46 | — | — | — | — | — | |
| 22B emaModel backbone=22B, Fine-tuned resolution=560px, Polyak averaging (EMA)=true, Training setup=Fine-tuned on ImageNet2023.02 | 95.05 | — | — | — | — | — | |
| APMP (Pre-trained on clean ImageNet)=false, Backbone=VIT-H/142024.10 | 94.9 | — | — | — | — | — | |
| LiTTraining Data Source=Private, Evaluation Protocol=Zero-shot2021.11 | 94.9 | — | — | — | — | — | |
| LiT ViT-gZero-shot=true2022.09 | 94.9 | — | — | — | — | — | |
| 22BModel backbone=22B, Fine-tuned resolution=560px, Polyak averaging (EMA)=false, Training setup=Fine-tuned on ImageNet2023.02 | 94.62 | — | — | — | — | — | |
| e/14 emaModel backbone=e/14, Fine-tuned resolution=560px, Polyak averaging (EMA)=true, Training setup=Fine-tuned on ImageNet2023.02 | 94.49 | — | — | — | — | — | |
| 22BModel backbone=22B, Fine-tuned resolution=384px, Polyak averaging (EMA)=false, Training setup=Fine-tuned on ImageNet2023.02 | 94.44 | — | — | — | — | — | |
| ViT-e/14Evaluation Protocol=Linear Probing (frozen), Resolution=224px2023.02 | 94.33 | — | — | — | — | — | |
| ViT-22BEvaluation Protocol=Linear Probing (frozen), Resolution=224px2023.02 | 94.27 | — | — | — | — | — | |
| G/14 emaModel backbone=G/14, Fine-tuned resolution=518px, Polyak averaging (EMA)=true, Training setup=Fine-tuned on ImageNet2023.02 | 94.22 | — | — | — | — | — | |
| LiTEvaluation Protocol=Zero-shot2022.05 | 93.9 | — | — | — | — | — | |
| OpenCLIP HScale=H2025.07 | 93.76 | — | — | — | — | — | |
| e/14 emaModel backbone=e/14, Fine-tuned resolution=384px, Polyak averaging (EMA)=true, Training setup=Fine-tuned on ImageNet2023.02 | 93.56 | — | — | — | — | — | |
| g/14 emaModel backbone=g/14, Fine-tuned resolution=518px, Polyak averaging (EMA)=true, Training setup=Fine-tuned on ImageNet2023.02 | 93.37 | — | — | — | — | — | |
| CoCa-BaseEvaluation Protocol=Zero-shot2022.05 | 93.2 | — | — | — | — | — | |
| OpenCLIP-VIT-H/14P (Pre-trained on clean ImageNet)=false, Backbone=VIT-H/142024.10 | 92.9 | — | — | — | — | — | |
| ViT-g/14Evaluation Protocol=Linear Probing (frozen), Resolution=224px2023.02 | 92.33 | — | — | — | — | — | |
| ViT-HParams (M)=632M, Labeled Data=zero-shot2025.05 | 92.3 | — | — | — | — | — | |
| ALIGNtransfer_mode=zero-shot, prompt_ensembling=true2021.02 | 92.2 | — | — | — | — | — | |
| ALIGNEvaluation Protocol=Zero-shot2022.05 | 92.2 | — | — | — | — | — | |
| ALIGNTraining Data Source=Private, Evaluation Protocol=Zero-shot2021.11 | 92.2 | — | — | — | — | — | |
| ALIGNZero-shot=true2021.11 | 92.2 | — | — | — | — | — | |
| ALIGNZero-shot=true2022.09 | 92.2 | — | — | — | — | — | |
| ViT-G/14Evaluation Protocol=Linear Probing (frozen), Resolution=224px2023.02 | 91.74 | — | — | — | — | — | |
| 22BModel backbone=22B, Training setup=JFT-only (zero-shot)2023.02 | 91.6 | — | — | — | — | — | |
| e/14Model backbone=e/14, Training setup=JFT-only (zero-shot)2023.02 | 90.6 | — | — | — | — | — | |
| L/16Model backbone=L/16, Fine-tuned resolution=384px, Polyak averaging (EMA)=false, Training setup=Fine-tuned on ImageNet2023.02 | 90.32 | — | — | — | — | — | |
| WiSE-FTModel=CLIP ViT-L/14@336px, Evaluation Protocol=End-to-End, Mixing Coefficient (alpha)=optimal2021.09 | 90.3 | — | — | — | — | — | |
| G/14Model backbone=G/14, Training setup=JFT-only (zero-shot)2023.02 | 90.2 | — | — | — | — | — | |
| ViT-L/16Evaluation Protocol=Linear Probing (frozen), Resolution=224px2023.02 | 89.92 | — | — | — | — | — | |
| WiSE-FTModel=CLIP ViT-L/14@336px, Evaluation Protocol=Linear Classifier, Mixing Coefficient (alpha)=optimal2021.09 | 89.8 | — | — | — | — | — | |
| g/14Model backbone=g/14, Training setup=JFT-only (zero-shot)2023.02 | 89.8 | — | — | — | — | — | |
| InternViT-6BParameters=5.9B, Evaluation Protocol=Linear Probing2023.12 | 89.8 | — | — | — | — | — | |
| WiSE-FTModel=CLIP ViT-L/14@336px, Evaluation Protocol=Linear Classifier, Mixing Coefficient (alpha)=0.52021.09 | 89.6 | — | — | — | — | — | |
| MobileCLIP-BEvaluation Protocol=zero-shot2023.11 | 89.6 | — | — | — | — | — | |
| WiSE-FTModel=CLIP ViT-L/14@336px, Evaluation Protocol=End-to-End, Mixing Coefficient (alpha)=0.52021.09 | 89.4 | — | — | — | — | — | |
| Zero-shotModel=CLIP ViT-L/14@336px, Evaluation Protocol=Zero-shot, Source=PyTorch2021.09 | 89 | — | — | — | — | — | |
| CLIPEvaluation Protocol=Zero-Shot2021.02 | 88.9 | — | — | — | — | — | |
| CLIPtransfer_mode=zero-shot, prompt_ensembling=true2021.02 | 88.9 | — | — | — | — | — | |
| CLIPEvaluation Protocol=Zero-shot2022.05 | 88.9 | — | — | — | — | — | |
| Zero-shotModel=CLIP ViT-L/14@336px, Evaluation Protocol=Zero-shot, Source=[82]2021.09 | 88.9 | — | — | — | — | — | |
| CLIPTraining Data Source=Private, Evaluation Protocol=Zero-shot2021.11 | 88.9 | — | — | — | — | — | |
| CLIPZero-shot=true2021.11 | 88.9 | — | — | — | — | — | |
| CLIPZero-shot=true2022.09 | 88.9 | — | — | — | — | — | |
| CLIPLanguage=English, Zero-shot=true, Backbone=ViT-L2022.11 | 87.9 | — | — | — | — | — | |
| AltCLIPTLanguage=English, Zero-shot=true, Backbone=ViT-L2022.11 | 87.9 | — | — | — | — | — | |
| OpenCLIP-GParameters=1.8B, Evaluation Protocol=Linear Probing2023.12 | 87.8 | — | — | — | — | — | |
| OpenCLIPArchitecture=ViT-G/14, Pretraining Data=LAION-2B, Resolution=224, Protocol=linear probe on frozen features2023.04 | 87.8 | — | — | — | — | — | |
| EVA-01-CLIP-gParameters=1.1B, Evaluation Protocol=Linear Probing2023.12 | 87.7 | — | — | — | — | — | |
| ViT-22BParameters=21.7B, Evaluation Protocol=Linear Probing, Training Data=JFT-3B2023.12 | 87.4 | — | — | — | — | — | |
| CLIPVision Encoder=ViT-L/14, zero-shot evaluation=true2022.12 | 87.32 | — | — | — | — | — | |
| AltCLIPLanguage=English, Zero-shot=true, Backbone=ViT-L2022.11 | 87.2 | — | — | — | — | — | |
| APMP (Pre-trained on clean ImageNet)=false, Backbone=VIT-L/142024.10 | 87.1 | — | — | — | — | — | |
| MobileCLIP-S2Evaluation Protocol=zero-shot2023.11 | 87 | — | — | — | — | — | |
| L/16Model backbone=L/16, Training setup=JFT-only (zero-shot)2023.02 | 86.8 | — | — | — | — | — | |
| ViT-L/14Params (M)=304M, Labeled Data=zero-shot2025.05 | 86.5 | — | — | — | — | — | |
| CLIP VIT-L/14P (Pre-trained on clean ImageNet)=false, Backbone=VIT-L/142024.10 | 85.9 | — | — | — | — | — | |
| CLIP + PACLVision Encoder=ViT-L/14, zero-shot evaluation=true2022.12 | 85.6 | — | — | — | — | — | |
| Fine-tuned LCModel=CLIP ViT-L/14@336px, Evaluation Protocol=Linear Classifier, Implementation=ours2021.09 | 85.3 | — | — | — | — | — | |
| MobileCLIP-S1Evaluation Protocol=zero-shot2023.11 | 84.7 | — | — | — | — | — | |
| GPT-4o2025.07 | 84.38 | — | — | — | — | — | |
| CLIPEvaluation Protocol=Linear Probe2021.02 | 84.2 | — | — | — | — | — | |
| Fine-tuned LCModel=CLIP ViT-L/14@336px, Evaluation Protocol=Linear Classifier, Source=[82]2021.09 | 84.2 | — | — | — | — | — | |
| ITOPre-training Dataset=DataComp-1B, Backbone=ViT-B/16, Training Epochs=10 epochs, Zero-shot protocol=EVA-CLIP2026.03 | 83.2 | — | — | — | — | — | |
| PaLI-XShots=0-shot, Resolution=2242023.05 | 82.96 | — | — | — | — | — | |
| B/16Model backbone=B/16, Fine-tuned resolution=384px, Polyak averaging (EMA)=false, Training setup=Fine-tuned on ImageNet2023.02 | 82.91 | — | — | — | — | — | |
| ITOBackbone=ViT-L/16, Pre-training Dataset=DataComp-1B, Training Epochs=1, Evaluation Protocol=Zero-shot2026.03 | 82.9 | — | — | — | — | — | |
| ITO sub2Pre-training Dataset=DataComp-1B, Backbone=ViT-B/16, Training Epochs=10 epochs, Zero-shot protocol=EVA-CLIP2026.03 | 82.9 | — | — | — | — | — | |
| DHOStudent Model=ViT-L/14, Params (M)=304M, Labeled Data=10%, Teacher Model=ViT-H/142025.05 | 82.8 | — | — | — | — | — | |
| AltCLIPLanguage=Chinese, Zero-shot=true, Backbone=ViT-L2022.11 | 82.5 | — | — | — | — | — | |
| ViT-B/16Evaluation Protocol=Linear Probing (frozen), Resolution=224px2023.02 | 82.5 | — | — | — | — | — | |
| ZERO+EnsembleBackbone=CLIP-ViT-B-16, Pre-trained=LAION-5B (2B English subset)2024.05 | 82.42 | — | — | — | — | — | |
| AltCLIPTLanguage=Chinese, Zero-shot=true, Backbone=ViT-L2022.11 | 82.1 | — | — | — | — | — | |
| Gemini 2.0 Flash2025.07 | 82.05 | — | — | — | — | — | |
| PaLIParameters=17B, Shots=0-shot, Resolution=2242023.05 | 81.97 | — | — | — | — | — | |
| M-CLIPLanguage=English, Zero-shot=true, Backbone=ViT-L2022.11 | 81.7 | — | — | — | — | — | |
| B/16Model backbone=B/16, Training setup=JFT-only (zero-shot)2023.02 | 81.4 | — | — | — | — | — | |
| CLIPPre-training Dataset=DataComp-1B, Backbone=ViT-B/16, Training Epochs=10 epochs, Zero-shot protocol=EVA-CLIP2026.03 | 81.2 | — | — | — | — | — | |
| PaLIParameters=3B, Resolution=224, Setting=Fine-tuning2023.05 | 81.11 | — | — | — | — | — | |
| ZEROBackbone=CLIP-ViT-B-16, Pre-trained=LAION-5B (2B English subset)2024.05 | 80.59 | — | — | — | — | — | |
| EnsembleBackbone=CLIP-ViT-B-16, Pre-trained=LAION-5B (2B English subset)2024.05 | 80.41 | — | — | — | — | — | |
| TPTBackbone=CLIP-ViT-B-16, Pre-trained=LAION-5B (2B English subset)2024.05 | 80.4 | — | — | — | — | — | |
| ITO sub2Backbone=ViT-L/16, Pre-training Dataset=DataComp-1B, Training Epochs=1, Evaluation Protocol=Zero-shot2026.03 | 80.4 | — | — | — | — | — |