Image Classification on ImageNet A
94.47Top-1 AccBest model on each test set (oracle)
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| Best model on each test set (oracle)Backbone=BASIC-L, Evaluation Protocol=fine-tuned, Model Selection Strategy=oracle (best on test set)2022.03 | 94.47 | — | — | — | — | — | |
| Greedy soupBackbone=BASIC-L, Evaluation Protocol=fine-tuned, Model Selection Strategy=greedy soup2022.03 | 94.17 | — | — | — | — | — | |
| Greedy ensembleBackbone=BASIC-L, Evaluation Protocol=fine-tuned, Model Selection Strategy=greedy ensemble2022.03 | 94.05 | — | — | — | — | — | |
| Best model on held out val setBackbone=BASIC-L, Evaluation Protocol=fine-tuned, Model Selection Strategy=best on val set2022.03 | 93.13 | — | — | — | — | — | |
| ViT-G/14 greedy soupBackbone=ViT/G-14, Model Selection Strategy=greedy soup2022.03 | 92.67 | — | — | — | — | — | |
| PEcore GSize=G2026.02 | 92.6 | — | — | — | — | — | |
| PEcore G (image only)Size=G, Input=image only2026.02 | 91.2 | — | — | — | — | — | |
| SigLIP2-g-optSize=g2026.02 | 90.5 | — | — | — | — | — | |
| CoCaEvaluation Protocol=Zero-shot2022.05 | 90.2 | — | — | — | — | — | |
| CoCaBackbone=CoCa, Evaluation Protocol=zero-shot2022.03 | 90.2 | — | — | — | — | — | |
| CoCaZero-shot=true2022.09 | 90.2 | — | — | — | — | — | |
| LiT-22B2026.02 | 90.1 | — | — | — | — | — | |
| 22B emaModel backbone=22B, Fine-tuned resolution=560px, Polyak averaging (EMA)=true, Training setup=Fine-tuned on ImageNet2023.02 | 89.12 | — | — | — | — | — | |
| PEcore LSize=L2026.02 | 89 | — | — | — | — | — | |
| EVA 18B+Size=18B, Version=plus2026.02 | 88.9 | — | — | — | — | — | |
| 22BModel backbone=22B, Fine-tuned resolution=560px, Polyak averaging (EMA)=false, Training setup=Fine-tuned on ImageNet2023.02 | 88.55 | — | — | — | — | — | |
| e/14 emaModel backbone=e/14, Fine-tuned resolution=560px, Polyak averaging (EMA)=true, Training setup=Fine-tuned on ImageNet2023.02 | 88.44 | — | — | — | — | — | |
| LiT ViT-eZero-shot=true2022.09 | 88 | — | — | — | — | — | |
| 22BModel backbone=22B, Fine-tuned resolution=384px, Polyak averaging (EMA)=false, Training setup=Fine-tuned on ImageNet2023.02 | 87.95 | — | — | — | — | — | |
| EVA-CLIP-18B#Param=17.5B, Zero-shot=true2025.05 | 87.3 | — | — | — | — | — | |
| EVA 18BSize=18B2026.02 | 87.3 | — | — | — | — | — | |
| e/14 emaModel backbone=e/14, Fine-tuned resolution=384px, Polyak averaging (EMA)=true, Training setup=Fine-tuned on ImageNet2023.02 | 87.16 | — | — | — | — | — | |
| G/14 emaModel backbone=G/14, Fine-tuned resolution=518px, Polyak averaging (EMA)=true, Training setup=Fine-tuned on ImageNet2023.02 | 86.95 | — | — | — | — | — | |
| MetaCLIP+ViSE2026.02 | 86.641 | — | — | — | — | — | |
| Only ViSE2026.02 | 86.641 | — | — | — | — | — | |
| LionBackbone=BASIC-L, Evaluation Protocol=Zero-shot2023.02 | 86.4 | — | — | — | — | — | |
| KD 2B to ViT-H, M+V+L4Backbone=ViT-H, Distillation=KD 2B, Target=M+V+L42026.02 | 86.26 | — | — | — | — | — | |
| g/14 emaModel backbone=g/14, Fine-tuned resolution=518px, Polyak averaging (EMA)=true, Training setup=Fine-tuned on ImageNet2023.02 | 86.12 | — | — | — | — | — | |
| CoCa-LargeEvaluation Protocol=Zero-shot2022.05 | 85.7 | — | — | — | — | — | |
| BASICEvaluation Protocol=Zero-shot2022.05 | 85.6 | — | — | — | — | — | |
| BASIC-LBackbone=BASIC-L, Evaluation Protocol=zero-shot2022.03 | 85.6 | — | — | — | — | — | |
| BASICZero-shot=true2021.11 | 85.6 | — | — | — | — | — | |
| AdafactorBackbone=BASIC-L, Evaluation Protocol=Zero-shot2023.02 | 85.6 | — | — | — | — | — | |
| BASICZero-shot=true2022.09 | 85.6 | — | — | — | — | — | |
| Seed-ViT#Param=532M, Zero-shot=true2025.05 | 85.5 | — | — | — | — | — | |
| NS EfficientNet-L22021.02 | 84.9 | — | — | — | — | — | |
| SigLIP2-L/16Backbone=SigLIP2-L, Patch Size=162026.02 | 84.3 | — | — | — | — | — | |
| APMP (Pre-trained on clean ImageNet)=false, Backbone=VIT-H/142024.10 | 84.2 | — | — | — | — | — | |
| KD 2B to ViT-H, ViseBackbone=ViT-H, Distillation=KD 2B, Target=Vise2026.02 | 83.98 | — | — | — | — | — | |
| ViT-22BEvaluation Protocol=Linear Probing (frozen), Resolution=224px2023.02 | 83.8 | — | — | — | — | — | |
| ViT-22BParameters=21.7B, Evaluation Protocol=Linear Probing, Training Data=JFT-3B2023.12 | 83.8 | — | — | — | — | — | |
| InternVL-C#Param=6B, Zero-shot=true2025.05 | 83.8 | — | — | — | — | — | |
| DenseModel=InternVL-C 6B, Pruning ratio=0%, Evaluation protocol=Zero-shot2026.02 | 83.8 | — | — | — | — | — | |
| InternVL-C2026.02 | 83.8 | — | — | — | — | — | |
| Noisy Student Trainingunlabeled data=300M images, backbone=EfficientNet2019.11 | 83.7 | — | — | — | — | — | |
| dino.txtResolution=336, Pre-training Dataset=LVTD-2.3B, Evaluation Protocol=Zero-shot2024.12 | 83.2 | — | — | — | — | — | |
| EVA-02-CLIP-L/14+zero-shot=true2023.03 | 82.9 | — | — | — | — | — | |
| EVA-02-CLIPResolution=336, Pre-training Dataset=Merged-2B, Evaluation Protocol=Zero-shot2024.12 | 82.9 | — | — | — | — | — | |
| EVA-02-CLIP-E/14+zero-shot=true2023.03 | 82.1 | — | — | — | — | — | |
| LiTTraining Data Source=Private, Evaluation Protocol=Zero-shot2021.11 | 81.8 | — | — | — | — | — | |
| LiT ViT-gZero-shot=true2022.09 | 81.8 | — | — | — | — | — | |
| ViT-e/14Evaluation Protocol=Linear Probing (frozen), Resolution=224px2023.02 | 81.56 | — | — | — | — | — | |
| WiSE-FTModel=CLIP ViT-L/14@336px, Evaluation Protocol=End-to-End, Mixing Coefficient (alpha)=optimal2021.09 | 81 | — | — | — | — | — | |
| EVA-02-CLIP-E/14zero-shot=true2023.03 | 80.4 | — | — | — | — | — | |
| dino.txtResolution=224, Pre-training Dataset=LVTD-2.3B, Evaluation Protocol=Zero-shot2024.12 | 80 | — | — | — | — | — | |
| WiSE-FTModel=CLIP ViT-L/14@336px, Evaluation Protocol=End-to-End, Mixing Coefficient (alpha)=0.52021.09 | 79.9 | — | — | — | — | — | |
| 22BModel backbone=22B, Training setup=JFT-only (zero-shot)2023.02 | 79.9 | — | — | — | — | — | |
| WiSE-FTModel=CLIP ViT-L/14@336px, Evaluation Protocol=Linear Classifier, Mixing Coefficient (alpha)=0.52021.09 | 79.7 | — | — | — | — | — | |
| WiSE-FTModel=CLIP ViT-L/14@336px, Evaluation Protocol=Linear Classifier, Mixing Coefficient (alpha)=optimal2021.09 | 79.7 | — | — | — | — | — | |
| DFN-5B-CLIP-H/14++#Param=632M, Zero-shot=true2025.05 | 79.6 | — | — | — | — | — | |
| DFN-H+2026.02 | 79.6 | — | — | — | — | — | |
| CAFormer-B36Params (M)=99, MACS (G)=72.2, Resolution=384, Pre-training=ImageNet-21K2022.10 | 79.5 | — | — | — | — | — | |
| LiTEvaluation Protocol=Zero-shot2022.05 | 79.4 | — | — | — | — | — | |
| OpenCLIP-VIT-H/14P (Pre-trained on clean ImageNet)=false, Backbone=VIT-H/142024.10 | 79.1 | — | — | — | — | — | |
| ViT-G/14Evaluation Protocol=Linear Probing (frozen), Resolution=224px2023.02 | 78.79 | — | — | — | — | — | |
| L/16Model backbone=L/16, Fine-tuned resolution=384px, Polyak averaging (EMA)=false, Training setup=Fine-tuned on ImageNet2023.02 | 78.65 | — | — | — | — | — | |
| Zero-shotModel=CLIP ViT-L/14@336px, Evaluation Protocol=Zero-shot, Source=PyTorch2021.09 | 77.7 | — | — | — | — | — | |
| ViT-g/14Evaluation Protocol=Linear Probing (frozen), Resolution=224px2023.02 | 77.51 | — | — | — | — | — | |
| OpenAI CLIP-L/14+zero-shot=true2023.03 | 77.5 | — | — | — | — | — | |
| InternViT-6BParameters=5.9B, Evaluation Protocol=Linear Probing2023.12 | 77.5 | — | — | — | — | — | |
| CLIPResolution=336, Pre-training Dataset=WIT-400M, Evaluation Protocol=Zero-shot2024.12 | 77.5 | — | — | — | — | — | |
| ViT-HParams (M)=632M, Labeled Data=zero-shot2025.05 | 77.4 | — | — | — | — | — | |
| CLIPEvaluation Protocol=Zero-Shot2021.02 | 77.2 | — | — | — | — | — | |
| CLIPtransfer_mode=zero-shot, prompt_ensembling=true2021.02 | 77.2 | — | — | — | — | — | |
| CLIPEvaluation Protocol=Zero-shot2022.05 | 77.2 | — | — | — | — | — | |
| Zero-shotModel=CLIP ViT-L/14@336px, Evaluation Protocol=Zero-shot, Source=[82]2021.09 | 77.2 | — | — | — | — | — | |
| CLIPTraining Data Source=Private, Evaluation Protocol=Zero-shot2021.11 | 77.2 | — | — | — | — | — | |
| CLIPZero-shot=true2021.11 | 77.2 | — | — | — | — | — | |
| CLIPZero-shot=true2022.09 | 77.2 | — | — | — | — | — | |
| SigLIP-L/16Backbone=SigLIP-L, Patch Size=162026.02 | 76.5 | — | — | — | — | — | |
| CoCa-BaseEvaluation Protocol=Zero-shot2022.05 | 76.4 | — | — | — | — | — | |
| SigLIPResolution=384, Pre-training Dataset=WebLi, Evaluation Protocol=Zero-shot2024.12 | 76.4 | — | — | — | — | — | |
| Fine-tuned LCModel=CLIP ViT-L/14@336px, Evaluation Protocol=Linear Classifier, Implementation=ours2021.09 | 76.1 | — | — | — | — | — | |
| EVA-02-CLIP-L/14zero-shot=true2023.03 | 76.1 | — | — | — | — | — | |
| DINOv2-gParameters=1.1B, Evaluation Protocol=Linear Probing2023.12 | 75.9 | — | — | — | — | — | |
| DINOv2Architecture=ViT-g/14, Pretraining Data=LVD-142M, Resolution=224, Protocol=linear probe on frozen features2023.04 | 75.9 | — | — | — | — | — | |
| ALIGNtransfer_mode=zero-shot, prompt_ensembling=true2021.02 | 75.8 | — | — | — | — | — | |
| ALIGNEvaluation Protocol=Zero-shot2022.05 | 75.8 | — | — | — | — | — | |
| ALIGNTraining Data Source=Private, Evaluation Protocol=Zero-shot2021.11 | 75.8 | — | — | — | — | — | |
| ALIGNZero-shot=true2021.11 | 75.8 | — | — | — | — | — | |
| ALIGNZero-shot=true2022.09 | 75.8 | — | — | — | — | — | |
| CLIPEvaluation Protocol=Linear Probe2021.02 | 75.3 | — | — | — | — | — | |
| Fine-tuned LCModel=CLIP ViT-L/14@336px, Evaluation Protocol=Linear Classifier, Source=[82]2021.09 | 75.3 | — | — | — | — | — | |
| e/14Model backbone=e/14, Training setup=JFT-only (zero-shot)2023.02 | 75.3 | — | — | — | — | — | |
| G/14Model backbone=G/14, Training setup=JFT-only (zero-shot)2023.02 | 75.1 | — | — | — | — | — | |
| MVT + FTMLLM=MMICL, Backbone=ViT-L2025.12 | 75.1 | — | — | — | — | — | |
| TRACERBackbone=CLIP ViT-L/142026.05 | 74.87 | — | 6.65 | — | — | — | |
| ViT-L/14Params (M)=304M, Labeled Data=zero-shot2025.05 | 74.6 | — | — | — | — | — | |
| FlattenGPTModel=InternVL-C 6B, Pruning ratio=20%, Evaluation protocol=Zero-shot2026.02 | 74.6 | — | — | — | — | — | |
| EVA-01-CLIP-g/14+zero-shot=true2023.03 | 74.1 | — | — | — | — | — |