Image Classification on ImageNet Distribution Shifts Summary
86.54Avg Shifts ScoreBest model on each test set (oracle)
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Best model on each test set (oracle)Backbone=BASIC-L, Evaluation Protocol=fine-tuned, Model Selection Strategy=oracle (best on test set)2022.03 | 86.54 | — | — | — | |
| Greedy soupBackbone=BASIC-L, Evaluation Protocol=fine-tuned, Model Selection Strategy=greedy soup2022.03 | 86.4 | — | — | — | |
| Greedy ensembleBackbone=BASIC-L, Evaluation Protocol=fine-tuned, Model Selection Strategy=greedy ensemble2022.03 | 86.2 | — | — | — | |
| Best model on held out val setBackbone=BASIC-L, Evaluation Protocol=fine-tuned, Model Selection Strategy=best on val set2022.03 | 85.63 | — | — | — | |
| CoCaBackbone=CoCa, Evaluation Protocol=zero-shot2022.03 | 85.54 | — | — | — | |
| ViT-G/14 greedy soupBackbone=ViT/G-14, Model Selection Strategy=greedy soup2022.03 | 85.02 | — | — | — | |
| BASIC-LBackbone=BASIC-L, Evaluation Protocol=zero-shot2022.03 | 84.06 | — | — | — | |
| WiSE-FTModel=CLIP ViT-L/14@336px, Evaluation Protocol=End-to-End, Mixing Coefficient (alpha)=optimal2021.09 | 77.4 | 81.9 | — | — | |
| WiSE-FTModel=CLIP ViT-L/14@336px, Evaluation Protocol=End-to-End, Mixing Coefficient (alpha)=0.52021.09 | 76.9 | 81.8 | — | — | |
| WiSE-FTModel=CLIP ViT-L/14@336px, Evaluation Protocol=Linear Classifier, Mixing Coefficient (alpha)=0.52021.09 | 75.9 | 79.8 | — | — | |
| WiSE-FTModel=CLIP ViT-L/14@336px, Evaluation Protocol=Linear Classifier, Mixing Coefficient (alpha)=optimal2021.09 | 75.9 | 80.2 | — | — | |
| Zero-shotModel=CLIP ViT-L/14@336px, Evaluation Protocol=Zero-shot, Source=PyTorch2021.09 | 73.4 | 75 | — | — | |
| Zero-shotModel=CLIP ViT-L/14@336px, Evaluation Protocol=Zero-shot, Source=[82]2021.09 | 73.3 | 74.8 | — | — | |
| Fine-tuned LCModel=CLIP ViT-L/14@336px, Evaluation Protocol=Linear Classifier, Implementation=ours2021.09 | 72.6 | 78.9 | — | — | |
| Fine-tuned LCModel=CLIP ViT-L/14@336px, Evaluation Protocol=Linear Classifier, Source=[82]2021.09 | 71.8 | 78.6 | — | — | |
| Fine-tuned E2EModel=CLIP ViT-L/14@336px, Evaluation Protocol=End-to-End, Implementation=ours2021.09 | 68.6 | 77.4 | — | — | |
| TPT + CoOpBackbone=ViT-B/162022.09 | 64.99 | — | 62.83 | — | |
| TPT + CoCoOpBackbone=ViT-B/162022.09 | 64.3 | — | 62.61 | — | |
| TPTBackbone=ViT-B/162022.09 | 62.44 | — | 60.81 | — | |
| CoCoOpBackbone=ViT-B/162022.09 | 62.13 | — | 59.91 | — | |
| CoOpBackbone=ViT-B/162022.09 | 61.72 | — | 59.28 | — | |
| Prompt EnsembleBackbone=ViT-B/162022.09 | 61.2 | — | 59.42 | — | |
| CLIPBackbone=ViT-B/162022.09 | 59.11 | — | 57.2 | — | |
| TPT + CoOpBackbone=ResNet-502022.09 | 49.55 | — | 45.75 | — | |
| TPT + CoCoOpBackbone=ResNet-502022.09 | 48.45 | — | 44.83 | — | |
| TPTBackbone=ResNet-502022.09 | 47.26 | — | 43.89 | — | |
| CoCoOpBackbone=ResNet-502022.09 | 46.81 | — | 42.82 | — | |
| CoOpBackbone=ResNet-502022.09 | 46.61 | — | 42.43 | — | |
| Prompt EnsembleBackbone=ResNet-502022.09 | 46.43 | — | 43.09 | — | |
| CLIPBackbone=ResNet-502022.09 | 44.18 | — | 40.69 | — | |
| DAREBackbone=CLIP ViT-B/16, Protocol=Linear Probing2026.01 | — | — | — | 58.45 | |
| DAREBackbone=CLIP ViT-B/16, Protocol=E2E-FT2026.01 | — | — | — | 53.38 | |
| E2E-FTBackbone=CLIP ViT-B/16, Protocol=E2E-FT2026.01 | — | — | — | 53.7 | |
| Linear ProbingBackbone=CLIP ViT-B/16, Protocol=Linear Probing2026.01 | — | — | — | 57.39 | |
| MERGETUNEBackbone=CLIP ViT-B/16, Protocol=Linear Probing2026.01 | — | — | — | 59.66 | |
| MERGETUNEBackbone=CLIP ViT-B/16, Protocol=E2E-FT2026.01 | — | — | — | 62.29 | |
| MERGETUNE + Weight ens.Backbone=CLIP ViT-B/16, Protocol=Linear Probing2026.01 | — | — | — | 60.23 | |
| MERGETUNE + Weight ens.Backbone=CLIP ViT-B/16, Protocol=E2E-FT2026.01 | — | — | — | 62.9 | |
| TIESBackbone=CLIP ViT-B/16, Protocol=Linear Probing2026.01 | — | — | — | 58.76 | |
| TIESBackbone=CLIP ViT-B/16, Protocol=E2E-FT2026.01 | — | — | — | 60.33 | |
| VRFBackbone=CLIP ViT-B/16, Protocol=Linear Probing2026.01 | — | — | — | 58.87 | |
| VRFBackbone=CLIP ViT-B/16, Protocol=E2E-FT2026.01 | — | — | — | 61.72 | |
| Weight ens.Backbone=CLIP ViT-B/16, Protocol=Linear Probing2026.01 | — | — | — | 58.56 | |
| Weight ens.Backbone=CLIP ViT-B/16, Protocol=E2E-FT2026.01 | — | — | — | 60.64 | |
| Zero-shot (CLIP)Backbone=CLIP ViT-B/16, Protocol=Zero-shot2026.01 | — | — | — | 58.42 |