Image Classification on ImageNet-Sketch
93.71Top-1 AccuracyKD 2B to ViT-H, M+V+L4
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| KD 2B to ViT-H, M+V+L4Backbone=ViT-H, Distillation=KD 2B, Target=M+V+L42026.02 | 93.71 | — | — | — | |
| KD 2B to ViT-H, ViseBackbone=ViT-H, Distillation=KD 2B, Target=Vise2026.02 | 85.04 | — | — | — | |
| PEcore GSize=G2026.02 | 83.7 | — | — | — | |
| PEcore G (image only)Size=G, Input=image only2026.02 | 82.7 | — | — | — | |
| SigLIP2-g-optSize=g2026.02 | 81 | — | — | — | |
| DFN-H+2026.02 | 80.5 | — | — | — | |
| PEcore LSize=L2026.02 | 80 | — | — | — | |
| EVA 18BSize=18B2026.02 | 78.8 | — | — | — | |
| SigLIP2-L/16Backbone=SigLIP2-L, Patch Size=162026.02 | 78.4 | — | — | — | |
| MetaCLIP+ViSE2026.02 | 77.764 | — | — | — | |
| Only ViSE2026.02 | 77.764 | — | — | — | |
| CoCaEvaluation Protocol=Zero-shot2022.05 | 77.6 | — | — | — | |
| CoCaBackbone=CoCa, Evaluation Protocol=zero-shot2022.03 | 77.6 | — | — | — | |
| Best model on each test set (oracle)Backbone=BASIC-L, Evaluation Protocol=fine-tuned, Model Selection Strategy=oracle (best on test set)2022.03 | 77.3 | — | — | — | |
| LionBackbone=BASIC-L, Evaluation Protocol=Zero-shot2023.02 | 77.2 | — | — | — | |
| Greedy soupBackbone=BASIC-L, Evaluation Protocol=fine-tuned, Model Selection Strategy=greedy soup2022.03 | 77.18 | — | — | — | |
| APMP (Pre-trained on clean ImageNet)=false, Backbone=VIT-H/142024.10 | 77.1 | — | — | — | |
| Best model on held out val setBackbone=BASIC-L, Evaluation Protocol=fine-tuned, Model Selection Strategy=best on val set2022.03 | 76.98 | — | — | — | |
| Greedy ensembleBackbone=BASIC-L, Evaluation Protocol=fine-tuned, Model Selection Strategy=greedy ensemble2022.03 | 76.63 | — | — | — | |
| InternVL-C2026.02 | 76.4 | — | — | — | |
| BASICEvaluation Protocol=Zero-shot2022.05 | 76.1 | — | — | — | |
| BASIC-LBackbone=BASIC-L, Evaluation Protocol=zero-shot2022.03 | 76.1 | — | — | — | |
| BASICZero-shot=true2021.11 | 76.1 | — | — | — | |
| AdafactorBackbone=BASIC-L, Evaluation Protocol=Zero-shot2023.02 | 76.1 | — | — | — | |
| CoCa-LargeEvaluation Protocol=Zero-shot2022.05 | 75.7 | — | — | — | |
| SigLIP 2Architecture=ViT-L/16, Pre-training Data=WebLI, Vision-Language Supervision=Yes, Evaluation Protocol=Linear Probing2025.07 | 75.5 | — | — | — | |
| EVA-CLIP-18BZero-shot=true2024.02 | 74.7 | — | — | — | |
| SigLIP-L/16Backbone=SigLIP-L, Patch Size=162026.02 | 74.4 | — | — | — | |
| InternVL-CZero-shot=true2024.02 | 74.3 | — | — | — | |
| EVA-CLIP-8BZero-shot=true2024.02 | 74.3 | — | — | — | |
| ViT-G/14 greedy soupBackbone=ViT/G-14, Model Selection Strategy=greedy soup2022.03 | 74.23 | — | — | — | |
| Model Soups ViT-GBackbone=ViT-G2025.07 | 74.23 | — | — | — | |
| SigLIPArchitecture=ViT-L/16, Pre-training Data=WebLI, Vision-Language Supervision=Yes, Evaluation Protocol=Linear Probing2025.07 | 73.6 | — | — | — | |
| PEcoreArchitecture=ViT-L/16, Pre-training Data=MC-2B, Vision-Language Supervision=Yes, Evaluation Protocol=Linear Probing2025.07 | 73.4 | — | — | — | |
| PaLI-XResolution=756, Setting=Fine-tuning2023.05 | 73.39 | — | — | — | |
| PaLI-XResolution=756, Setting=Fine-tuning, Steps=2.2x more steps2023.05 | 73.37 | — | — | — | |
| DFN5B-CLIP-H/14+Zero-shot=true2024.02 | 73.3 | — | — | — | |
| OpenCLIP HScale=H2025.07 | 73.24 | — | — | — | |
| OpenCLIP-VIT-H/14P (Pre-trained on clean ImageNet)=false, Backbone=VIT-H/142024.10 | 72.8 | — | — | — | |
| DFN5B-CLIP-H/14Zero-shot=true2024.02 | 72.8 | — | — | — | |
| PaLI-XResolution=224, Setting=Fine-tuning2023.05 | 72.56 | — | — | — | |
| EVA-02-CLIP-E/14+Zero-shot=true2024.02 | 72.2 | — | — | — | |
| ViT-HParams (M)=632M, Labeled Data=zero-shot2025.05 | 71.7 | — | — | — | |
| CoCa-BaseEvaluation Protocol=Zero-shot2022.05 | 71.7 | — | — | — | |
| PaLIParameters=17B, Resolution=224, Setting=Fine-tuning2023.05 | 71.21 | — | — | — | |
| PaLIParameters=3B, Resolution=224, Setting=Fine-tuning2023.05 | 70 | — | — | — | |
| OpenCLIP-G/14Zero-shot=true2024.02 | 69.9 | — | — | — | |
| Gemini 2.0 Flash2025.07 | 69.43 | — | — | — | |
| InternViT-6BParameters=5.9B, Evaluation Protocol=Linear Probing2023.12 | 69.1 | — | — | — | |
| SigLIP2-B/16Backbone=SigLIP2-B, Patch Size=162026.02 | 68.9 | — | — | — | |
| EVA-01-CLIP-g/14+Zero-shot=true2024.02 | 68.4 | — | — | — | |
| SigLIP-B/16Backbone=SigLIP-B, Patch Size=162026.02 | 67.9 | — | — | — | |
| EVA-01-CLIP-g/14Zero-shot=true2024.02 | 67.6 | — | — | — | |
| GPT-4o2025.07 | 67.3 | — | — | — | |
| OpenCLIP-GParameters=1.8B, Evaluation Protocol=Linear Probing2023.12 | 66.4 | — | — | — | |
| OpenCLIPArchitecture=ViT-G/14, Pretraining Data=LAION-2B, Resolution=224, Protocol=linear probe on frozen features2023.04 | 66.4 | — | — | — | |
| OpenCLIPArchitecture=ViT-G/14, Pre-training Data=LAION-2B, Vision-Language Supervision=Yes, Evaluation Protocol=Linear Probing2025.07 | 66.4 | — | — | — | |
| PEcore BSize=B2026.02 | 66.1 | — | — | — | |
| WiSE-FTModel=CLIP ViT-L/14@336px, Evaluation Protocol=End-to-End, Mixing Coefficient (alpha)=optimal2021.09 | 65 | — | — | — | |
| ALIGNEvaluation Protocol=Zero-shot2022.05 | 64.8 | — | — | — | |
| ALIGNZero-shot=true2021.11 | 64.8 | — | — | — | |
| WiSE-FTModel=CLIP ViT-L/14@336px, Evaluation Protocol=End-to-End, Mixing Coefficient (alpha)=0.52021.09 | 64.7 | — | — | — | |
| MobileCLIP-BEvaluation Protocol=zero-shot2023.11 | 64.5 | — | — | — | |
| PaLIParameters=17B, Shots=0-shot, Resolution=2242023.05 | 63.83 | — | — | — | |
| CLIP + PACLVision Encoder=ViT-L/14, zero-shot evaluation=true2022.12 | 63.23 | — | — | — | |
| EVA-01-CLIP-gParameters=1.1B, Evaluation Protocol=Linear Probing2023.12 | 63.1 | — | — | — | |
| WiSE-FTModel=CLIP ViT-L/14@336px, Evaluation Protocol=Linear Classifier, Mixing Coefficient (alpha)=0.52021.09 | 63 | — | — | — | |
| WiSE-FTModel=CLIP ViT-L/14@336px, Evaluation Protocol=Linear Classifier, Mixing Coefficient (alpha)=optimal2021.09 | 63 | — | — | — | |
| DINOv2-gParameters=1.1B, Evaluation Protocol=Linear Probing2023.12 | 62.5 | — | — | — | |
| DINOv2Architecture=ViT-g/14, Pretraining Data=LVD-142M, Resolution=224, Protocol=linear probe on frozen features2023.04 | 62.5 | — | — | — | |
| DINOv2Architecture=ViT-G/14, Pre-training Data=LVD-142M, Vision-Language Supervision=No, Evaluation Protocol=Linear Probing2025.07 | 62.5 | — | — | — | |
| APMP (Pre-trained on clean ImageNet)=false, Backbone=VIT-L/142024.10 | 62.2 | — | — | — | |
| MobileCLIP-S2Evaluation Protocol=zero-shot2023.11 | 62.2 | — | — | — | |
| DHOStudent Model=ViT-L/14, Params (M)=304M, Labeled Data=10%, Teacher Model=ViT-H/142025.05 | 61.7 | — | — | — | |
| PaLI-XShots=0-shot, Resolution=2242023.05 | 61.58 | — | — | — | |
| DHOStudent Model=ViT-L/14, Params (M)=304M, Labeled Data=1%, Teacher Model=ViT-H/142025.05 | 61.5 | — | — | — | |
| Zero-shotModel=CLIP ViT-L/14@336px, Evaluation Protocol=Zero-shot, Source=PyTorch2021.09 | 60.9 | — | — | — | |
| Web-SSLArchitecture=ViT-G/14, Pre-training Data=MC-2B, Vision-Language Supervision=No, Evaluation Protocol=Linear Probing2025.07 | 60.9 | — | — | — | |
| FrancaArchitecture=ViT-G/14, Pre-training Data=LAION-600M, Vision-Language Supervision=No, Evaluation Protocol=Linear Probing2025.07 | 60.6 | — | — | — | |
| MLCDPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Protocol=Zero-shot2024.07 | 60.4 | — | — | — | |
| MobileCLIP-S1Evaluation Protocol=zero-shot2023.11 | 60.3 | — | — | — | |
| CLIPEvaluation Protocol=Zero-Shot2021.02 | 60.2 | — | — | — | |
| CLIPEvaluation Protocol=Zero-shot2022.05 | 60.2 | — | — | — | |
| Zero-shotModel=CLIP ViT-L/14@336px, Evaluation Protocol=Zero-shot, Source=[82]2021.09 | 60.2 | — | — | — | |
| CLIPModel=ViT-L/14@336, Pre-training Data=WIT-400M, Image size=3362022.12 | 60.2 | — | — | — | |
| CLIPZero-shot=true2021.11 | 60.2 | — | — | — | |
| FLIPModel=ViT-L/14, Pre-training Data=LAION-400M, Image size=2242022.12 | 59.9 | — | — | — | |
| FLIPPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Protocol=Zero-shot2024.07 | 59.9 | — | — | — | |
| CLIPVision Encoder=ViT-L/14, zero-shot evaluation=true2022.12 | 59.71 | — | — | — | |
| CLIPModel=ViT-L/14@336, Pre-training Data=WIT-400M, Evaluation Protocol=our eval., Image size=3362022.12 | 59.7 | — | — | — | |
| CLIPLanguage=English, Zero-shot=true, Backbone=ViT-L2022.11 | 59.6 | — | — | — | |
| DINOv2Architecture=ViT-L/14, Pretraining Data=LVD-142M, Resolution=224, Protocol=linear probe on frozen features2023.04 | 59.3 | — | — | — | |
| DINOv2Architecture=ViT-L/14, Pre-training Data=LVD-142M, Vision-Language Supervision=No, Evaluation Protocol=Linear Probing, Note=distilled from DINOv2-G on LVD-142M2025.07 | 59.3 | — | — | — | |
| ViT-L/14Params (M)=304M, Labeled Data=zero-shot2025.05 | 59.2 | — | — | — | |
| AltCLIPTLanguage=English, Zero-shot=true, Backbone=ViT-L2022.11 | 59.2 | — | — | — | |
| FixResNeXt101-32x48d V22021.02 | 59.1 | — | — | — | |
| CLIP VIT-L/14P (Pre-trained on clean ImageNet)=false, Backbone=VIT-L/142024.10 | 58.8 | — | — | — | |
| Fine-tuned LCModel=CLIP ViT-L/14@336px, Evaluation Protocol=Linear Classifier, Implementation=ours2021.09 | 58.7 | — | — | — | |
| AltCLIPLanguage=English, Zero-shot=true, Backbone=ViT-L2022.11 | 58.7 | — | — | — | |
| CLIPModel=ViT-L/14, Pre-training Data=LAION-400M, Evaluation Protocol=our repro., Image size=2242022.12 | 58.7 | — | — | — |