Image Classification on ObjectNet
91.6AccuracyDFN-H+
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| DFN-H+2026.02 | 91.6 | — | — | — | — | — | — | |
| SigLIP2-g-optSize=g2026.02 | 91.5 | — | — | — | — | — | — | |
| PEcore GSize=G2026.02 | 91.4 | — | — | — | — | — | — | |
| PEcore G (image only)Size=G, Input=image only2026.02 | 91 | — | — | — | — | — | — | |
| SigLIP2-L/16Backbone=SigLIP2-L, Patch Size=162026.02 | 90 | — | — | — | — | — | — | |
| SigLIP-L/16Backbone=SigLIP-L, Patch Size=162026.02 | 89.4 | — | — | — | — | — | — | |
| LiT-22B2026.02 | 87.6 | — | — | — | — | — | — | |
| PEcore LSize=L2026.02 | 87.2 | — | — | — | — | — | — | |
| EVA 18BSize=18B2026.02 | 86 | — | — | — | — | — | — | |
| InternVL-C2026.02 | 85.8 | — | — | — | — | — | — | |
| XrayViT=EViT-2B, Backbone Type=Ours, Resolution=336, Active Tokens=2882026.02 | 83.8 | — | — | — | — | — | — | |
| LionBackbone=BASIC-L, Evaluation Protocol=Zero-shot2023.02 | 82.9 | — | — | — | — | — | — | |
| CoCaBackbone=CoCa, Evaluation Protocol=zero-shot2022.03 | 82.7 | — | — | — | — | — | — | |
| LiTTraining Data Source=Private, Evaluation Protocol=Zero-shot2021.11 | 82.5 | — | — | — | — | — | — | |
| BASIC-LBackbone=BASIC-L, Evaluation Protocol=zero-shot2022.03 | 82.3 | — | — | — | — | — | — | |
| AdafactorBackbone=BASIC-L, Evaluation Protocol=Zero-shot2023.02 | 82.3 | — | — | — | — | — | — | |
| MetaCLIP+ViSE2026.02 | 82.221 | — | — | — | — | — | — | |
| Only ViSE2026.02 | 82.221 | — | — | — | — | — | — | |
| EVA-CLIP-18B#Param=17.5B, Zero-shot=true2025.05 | 82.2 | — | — | — | — | — | — | |
| KD 2B to ViT-H, M+V+L4Backbone=ViT-H, Distillation=KD 2B, Target=M+V+L42026.02 | 81.63 | — | — | — | — | — | — | |
| Best model on each test set (oracle)Backbone=BASIC-L, Evaluation Protocol=fine-tuned, Model Selection Strategy=oracle (best on test set)2022.03 | 80.94 | — | — | — | — | — | — | |
| InternVL-C#Param=6B, Zero-shot=true2025.05 | 80.6 | — | — | — | — | — | — | |
| PEcoreViT=G/14, Backbone Type=Weakly-supervised backbones2026.02 | 80.2 | — | — | — | — | — | — | |
| Greedy soupBackbone=BASIC-L, Evaluation Protocol=fine-tuned, Model Selection Strategy=greedy soup2022.03 | 79.94 | — | — | — | — | — | — | |
| Greedy ensembleBackbone=BASIC-L, Evaluation Protocol=fine-tuned, Model Selection Strategy=greedy ensemble2022.03 | 79.91 | — | — | — | — | — | — | |
| Seed-ViT#Param=532M, Zero-shot=true2025.05 | 79.2 | — | — | — | — | — | — | |
| DINOv3ViT=7B/16, Backbone Type=Self-supervised backbones2026.02 | 79 | — | — | — | — | — | — | |
| SigLIP 2ViT=g/16, Backbone Type=Weakly-supervised backbones2026.02 | 78.6 | — | — | — | — | — | — | |
| ViT-G/14 greedy soupBackbone=ViT/G-14, Model Selection Strategy=greedy soup2022.03 | 78.52 | — | — | — | — | — | — | |
| Best model on held out val setBackbone=BASIC-L, Evaluation Protocol=fine-tuned, Model Selection Strategy=best on val set2022.03 | 78.09 | — | — | — | — | — | — | |
| DFN-5B-CLIP-H/14++#Param=632M, Zero-shot=true2025.05 | 78 | — | — | — | — | — | — | |
| KD 2B to ViT-H, ViseBackbone=ViT-H, Distillation=KD 2B, Target=Vise2026.02 | 76.91 | — | — | — | — | — | — | |
| Dehghani et al.ViT=22B/14, Backbone Type=Supervised backbones, Evaluation Protocol=Dehghani et al. (2023) protocol (*)2026.02 | 74.3 | — | — | — | — | — | — | |
| SigLIP2-B/16Backbone=SigLIP2-B, Patch Size=162026.02 | 73.6 | — | — | — | — | — | — | |
| OpenCLIP-G/14#Param=1.8B, Zero-shot=true2025.05 | 73 | — | — | — | — | — | — | |
| CLIP ViT L/14mode=zero-shot2022.01 | 72.3 | — | — | — | — | — | — | |
| CLIPTraining Data Source=Private, Evaluation Protocol=Zero-shot2021.11 | 72.3 | — | — | — | — | — | — | |
| WiSE-FTModel=CLIP ViT-L/14@336px, Evaluation Protocol=End-to-End, Mixing Coefficient (alpha)=optimal2021.09 | 72.1 | — | — | — | — | — | — | |
| PEcore BSize=B2026.02 | 71.9 | — | — | — | — | — | — | |
| EVA-CLIPViT=18B/14, Backbone Type=Weakly-supervised backbones2026.02 | 71.9 | — | — | — | — | — | — | |
| Chen et al.ViT=e/14, Backbone Type=Supervised backbones, Evaluation Protocol=Dehghani et al. (2023) protocol (*)2026.02 | 71.5 | — | — | — | — | — | — | |
| WiSE-FTModel=CLIP ViT-L/14@336px, Evaluation Protocol=End-to-End, Mixing Coefficient (alpha)=0.52021.09 | 71.1 | — | — | — | — | — | — | |
| WiSE-FTModel=CLIP ViT-L/14@336px, Evaluation Protocol=Linear Classifier, Mixing Coefficient (alpha)=0.52021.09 | 70.7 | — | — | — | — | — | — | |
| WiSE-FTModel=CLIP ViT-L/14@336px, Evaluation Protocol=Linear Classifier, Mixing Coefficient (alpha)=optimal2021.09 | 70.7 | — | — | — | — | — | — | |
| SigLIP-B/16Backbone=SigLIP-B, Patch Size=162026.02 | 70.7 | — | — | — | — | — | — | |
| ViT/G-14Backbone=ViT/G-142022.03 | 70.53 | — | — | — | — | — | — | |
| Zero-shotModel=CLIP ViT-L/14@336px, Evaluation Protocol=Zero-shot, Source=[82]2021.09 | 70 | — | — | — | — | — | — | |
| TRACERBackbone=CLIP ViT-L/142026.05 | 69.76 | — | — | — | — | — | 0.0918 | |
| Web-DINOViT=7B/14, Backbone Type=Self-supervised backbones2026.02 | 69.7 | — | — | — | — | — | — | |
| Zhai et al.ViT=G/14, Backbone Type=Supervised backbones, Evaluation Protocol=Dehghani et al. (2023) protocol (*)2026.02 | 69.6 | — | — | — | — | — | — | |
| MobileCLIP-BEvaluation Protocol=zero-shot2023.11 | 69.4 | — | — | — | — | — | — | |
| Zero-shotModel=CLIP ViT-L/14@336px, Evaluation Protocol=Zero-shot, Source=PyTorch2021.09 | 69.1 | — | — | — | — | — | — | |
| AIMv2ViT=3B/14, Backbone Type=Weakly-supervised backbones2026.02 | 69 | — | — | — | — | — | — | |
| CLIPPre-training Data=WIT-400M, Backbone=ViT-L/14, Evaluation Protocol=Zero-shot2024.07 | 68.6 | — | — | — | — | — | — | |
| AM-RADIOv2.5ViT=g/14, Backbone Type=Agglomerative backbones2026.02 | 68.4 | — | — | — | — | — | — | |
| CaRotBackbone=CLIP ViT-L/142026.05 | 68.05 | — | — | — | — | — | 0.1051 | |
| TuneCLIPBase Model=SigLIP ViT-B/162026.01 | 67.93 | — | — | — | — | — | — | |
| Fine-tuned LCModel=CLIP ViT-L/14@336px, Evaluation Protocol=Linear Classifier, Implementation=ours2021.09 | 67.2 | — | — | — | — | — | — | |
| MobileCLIP-S2Evaluation Protocol=zero-shot2023.11 | 66.6 | — | — | — | — | — | — | |
| ZSBackbone=CLIP ViT-L/14, Protocol=Zero-Shot2026.05 | 66.59 | — | — | — | — | — | 0.0852 | |
| DINOv2ViT=g/14, Backbone Type=Self-supervised backbones2026.02 | 66.4 | — | — | — | — | — | — | |
| Fine-tuned LCModel=CLIP ViT-L/14@336px, Evaluation Protocol=Linear Classifier, Source=[82]2021.09 | 66.2 | — | — | — | — | — | — | |
| FLYPBackbone=CLIP ViT-L/142026.05 | 66.15 | — | — | — | — | — | 0.1903 | |
| RegNetY 128GFmode=zero-shot, Platt scaling=true2022.01 | 64.3 | — | — | — | — | — | — | |
| MobileCLIP-S1Evaluation Protocol=zero-shot2023.11 | 63.4 | — | — | — | — | — | — | |
| Fine-tuned E2EModel=CLIP ViT-L/14@336px, Evaluation Protocol=End-to-End, Implementation=ours2021.09 | 63.3 | — | — | — | — | — | — | |
| MLCDPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Protocol=Zero-shot2024.07 | 62.7 | — | — | — | — | — | — | |
| LP-FTBackbone=CLIP ViT-L/14, Protocol=Linear Probing then Fine-Tuning2026.05 | 60.12 | — | — | — | — | — | 0.2572 | |
| ViT H/14mode=zero-shot, Platt scaling=true2022.01 | 60 | — | — | — | — | — | — | |
| FTBackbone=CLIP ViT-L/14, Protocol=Fine-Tuning2026.05 | 59.76 | — | — | — | — | — | 0.2865 | |
| RegNetY 32GFmode=zero-shot, Platt scaling=false2022.01 | 59.1 | — | — | — | — | — | — | |
| FLIPPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Protocol=Zero-shot2024.07 | 59.1 | — | — | — | — | — | — | |
| OpenCLIPPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Protocol=Zero-shot2024.07 | 58.8 | — | — | — | — | — | — | |
| ViT L/16mode=zero-shot, Platt scaling=true2022.01 | 57.3 | — | — | — | — | — | — | |
| TuneCLIPBase Model=OpenAI ViT-B/162026.01 | 57.08 | — | — | — | — | — | — | |
| MERGETUNE + Weight ens.Backbone=CLIP ViT-B/16, Protocol=E2E-FT2026.01 | 56.84 | — | — | — | — | — | — | |
| OpenCLIPBase Model=OpenAI ViT-B/162026.01 | 56.67 | — | — | — | — | — | — | |
| VRFBackbone=CLIP ViT-B/16, Protocol=E2E-FT2026.01 | 56.41 | — | — | — | — | — | — | |
| MERGETUNEBackbone=CLIP ViT-B/16, Protocol=E2E-FT2026.01 | 56.22 | — | — | — | — | — | — | |
| MobileCLIP-S0Evaluation Protocol=zero-shot2023.11 | 55.9 | — | — | — | — | — | — | |
| Weight ens.Backbone=CLIP ViT-B/16, Protocol=E2E-FT2026.01 | 55.71 | — | — | — | — | — | — | |
| FastCLIPBase Model=OpenAI ViT-B/162026.01 | 55.37 | — | — | — | — | — | — | |
| BaselineBase Model=OpenAI ViT-B/162026.01 | 55.31 | — | — | — | — | — | — | |
| BaselineBase Model=SigLIP ViT-B/162026.01 | 55.09 | — | — | — | — | — | — | |
| TIESBackbone=CLIP ViT-B/16, Protocol=E2E-FT2026.01 | 54.87 | — | — | — | — | — | — | |
| RegNetY 16GFmode=zero-shot, Platt scaling=false2022.01 | 54.8 | — | — | — | — | — | — | |
| LiTTraining Data Source=Public, Evaluation Protocol=Zero-shot2021.11 | 54.5 | — | — | — | — | — | — | |
| FrancaViT=g/14, Backbone Type=Self-supervised backbones2026.02 | 54.5 | — | — | — | — | — | — | |
| Zero-shot (CLIP)Backbone=CLIP ViT-B/16, Protocol=Zero-shot2026.01 | 54.23 | — | — | — | — | — | — | |
| RegNetY 128GFmode=zero-shot, Platt scaling=false2022.01 | 54.2 | — | — | — | — | — | — | |
| MERGETUNE + Weight ens.Backbone=CLIP ViT-B/16, Protocol=Linear Probing2026.01 | 54.01 | — | — | — | — | — | — | |
| HUVREncoder=ViT-L, Dimension=32, Compression=HUVR, Evaluation Protocol=Linear probing2026.01 | 53.9 | — | — | — | — | — | — | |
| MERGETUNEBackbone=CLIP ViT-B/16, Protocol=Linear Probing2026.01 | 53.43 | — | — | — | — | — | — | |
| VRFBackbone=CLIP ViT-B/16, Protocol=Linear Probing2026.01 | 53.36 | — | — | — | — | — | — | |
| TIESBackbone=CLIP ViT-B/16, Protocol=Linear Probing2026.01 | 53.13 | — | — | — | — | — | — | |
| Weight ens.Backbone=CLIP ViT-B/16, Protocol=Linear Probing2026.01 | 53.07 | — | — | — | — | — | — | |
| DAREBackbone=CLIP ViT-B/16, Protocol=Linear Probing2026.01 | 53.01 | — | — | — | — | — | — | |
| ViT H/14mode=zero-shot, Platt scaling=false2022.01 | 52.4 | — | — | — | — | — | — | |
| FTBackbone=ViT-L/14, Step=Step 12026.05 | 52.39 | — | — | — | — | — | — | |
| Linear ProbingBackbone=CLIP ViT-B/16, Protocol=Linear Probing2026.01 | 52.28 | — | — | — | — | — | — |