Image Classification on STL-10
99.2AccuracyCLIP* L/14
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| CLIP* L/14MAP head=true, Grain=Coarse, Frozen representation=true, backbone=ViT-L/142023.06 | 99.2 | — | — | — | — | — | |
| CapPa L/14MAP head=true, Grain=Coarse, Frozen representation=true, backbone=ViT-L/142023.06 | 99.1 | — | — | — | — | — | |
| IndividualBackbone=ViT-B/16, Pre-trained=ImageNet-21k, Number of Merged Models=12024.05 | 99.07 | — | — | — | — | — | |
| EsViTBackbone=Swin-T, Evaluation Protocol=Linear probe2021.06 | 98.9 | — | — | — | — | — | |
| CLIP* (16k)MAP head=true, Grain=Coarse, Frozen representation=true2023.06 | 98.5 | — | — | — | — | — | |
| EMR-MERGINGBackbone=ViT-B/16, Pre-trained=ImageNet-21k, Number of Merged Models=302024.05 | 98.41 | — | — | — | — | — | |
| TuneCLIPBase Model=SigLIP ViT-B/162026.01 | 98.37 | — | — | — | — | — | |
| OpenCLIP VIT-H/14zero-shot=true2023.03 | 98.3 | — | — | — | — | — | |
| TuneCLIPBase Model=OpenAI ViT-B/162026.01 | 98.26 | — | — | — | — | — | |
| BaselineBase Model=OpenAI ViT-B/162026.01 | 98.25 | — | — | — | — | — | |
| BaselineBase Model=SigLIP ViT-B/162026.01 | 98.21 | — | — | — | — | — | |
| LaCLIPArchitecture=ViT-B/16, Pre-training Data=LAION-400M, Protocol=5-way 5-shot2023.05 | 98.1 | — | — | — | — | — | |
| CapPaMAP head=true, Grain=Coarse, Frozen representation=true2023.06 | 98.1 | — | — | — | — | — | |
| CLIP* (8k)MAP head=true, Grain=Coarse, Frozen representation=true2023.06 | 98 | — | — | — | — | — | |
| SupervisedBackbone=Swin-T, Evaluation Protocol=Linear probe2021.06 | 97.9 | — | — | — | — | — | |
| OpenCLIPBase Model=OpenAI ViT-B/162026.01 | 97.87 | — | — | — | — | — | |
| CapMAP head=true, Grain=Coarse, Frozen representation=true2023.06 | 97.7 | — | — | — | — | — | |
| CLIPBackbone=ResNet-50, Evaluation Protocol=Linear probe2021.06 | 97.2 | — | — | — | — | — | |
| LaCLIPArchitecture=ViT-B/32, Pre-training Data=LAION-400M, Protocol=5-way 5-shot2023.05 | 97.2 | — | — | — | — | — | |
| TuneCLIPBase Model=OpenAI ViT-B/322026.01 | 97.2 | — | — | — | — | — | |
| BaselineBase Model=OpenAI ViT-B/322026.01 | 97.13 | — | — | — | — | — | |
| CLIPArchitecture=ViT-B/16, Pre-training Data=LAION-400M, Protocol=5-way 5-shot2023.05 | 97.1 | — | — | — | — | — | |
| SupervisedBackbone=ResNet-50, Evaluation Protocol=Linear probe, Reproduced=true2021.06 | 97 | — | — | — | — | — | |
| FastCLIPBase Model=OpenAI ViT-B/162026.01 | 96.58 | — | — | — | — | — | |
| BaselineBase Model=LAION ViT-B/322026.01 | 96.56 | — | — | — | — | — | |
| TuneCLIPBase Model=LAION ViT-B/322026.01 | 96.47 | — | — | — | — | — | |
| SupervisedBackbone=ResNet-50, Evaluation Protocol=Linear probe2021.06 | 96.4 | — | — | — | — | — | |
| OpenCLIPBase Model=OpenAI ViT-B/322026.01 | 96.36 | — | — | — | — | — | |
| CLIPArchitecture=ViT-B/32, Pre-training Data=LAION-400M, Protocol=5-way 5-shot2023.05 | 96.1 | — | — | — | — | — | |
| LaCLIPArchitecture=ViT-B/16, Pre-training Data=CC12M, Protocol=5-way 5-shot2023.05 | 95.9 | — | — | — | — | — | |
| UNCHABackbone=ViT-B/16, Zero-shot=true2026.03 | 95.7 | — | — | — | — | — | |
| LaCLIPArchitecture=ViT-B/16, Pre-training Data=RedCaps, Protocol=5-way 5-shot2023.05 | 95.6 | — | — | — | — | — | |
| Diffusion Classifierzero-shot=true2023.03 | 95.4 | — | — | — | — | — | |
| OpenCLIPBase Model=LAION ViT-B/322026.01 | 95.16 | — | — | — | — | — | |
| HyCoCLIPBackbone=ViT-B/16, Zero-shot=true2026.03 | 95 | — | — | — | — | — | |
| CLIPArchitecture=ViT-B/16, Pre-training Data=RedCaps, Protocol=5-way 5-shot2023.05 | 94.8 | — | — | — | — | — | |
| UNCHABackbone=ViT-S/16, Zero-shot=true2026.03 | 94.4 | — | — | — | — | — | |
| CLIP ResNet-50zero-shot=true2023.03 | 94.3 | — | — | — | — | — | |
| LaSLIPArchitecture=ViT-B/16, Pre-training Data=CC12M, Protocol=5-way 5-shot2023.05 | 94.2 | — | — | — | — | — | |
| Standard CLIPEvaluation Protocol=zero-shot2024.11 | 94 | — | — | — | — | — | |
| B-cosified RN-50 CLIPPre-training Dataset=ImageNet [16], Learning Scheduler=Cosine, Evaluation Protocol=zero-shot2024.11 | 94 | — | — | — | — | — | |
| B-cosified RN-50 CLIPPre-training Dataset=ImageNet [16], Learning Scheduler=Cyclic, Evaluation Protocol=zero-shot2024.11 | 94 | — | — | — | — | — | |
| OpenCLIPBase Model=SigLIP ViT-B/162026.01 | 93.81 | — | — | — | — | — | |
| B-cosified RN-50 CLIPPre-training Dataset=CC3M [47], Learning Scheduler=Cosine, Evaluation Protocol=zero-shot2024.11 | 93 | — | — | — | — | — | |
| B-cosified RN-50 CLIPPre-training Dataset=CC3M [47], Learning Scheduler=Cyclic, Evaluation Protocol=zero-shot2024.11 | 93 | — | — | — | — | — | |
| MERUBackbone=ViT-B/16, Zero-shot=true2026.03 | 92.8 | — | — | — | — | — | |
| FastCLIPBase Model=OpenAI ViT-B/322026.01 | 92.68 | — | — | — | — | — | |
| CLIPArchitecture=ViT-B/16, Pre-training Data=CC12M, Protocol=5-way 5-shot2023.05 | 92.6 | — | — | — | — | — | |
| SLIPArchitecture=ViT-B/16, Pre-training Data=CC12M, Protocol=5-way 5-shot2023.05 | 92.5 | — | — | — | — | — | |
| HyCoCLIPBackbone=ViT-S/16, Zero-shot=true2026.03 | 92.5 | — | — | — | — | — | |
| CLIPBackbone=ViT-B/16, Zero-shot=true2026.03 | 92.4 | — | — | — | — | — | |
| FastCLIPBase Model=LAION ViT-B/322026.01 | 91.73 | — | — | — | — | — | |
| Mixed Barlow TwinsBackbone=ResNet-50, Epochs=2000, Evaluation Protocol=linear2023.12 | 91.7 | — | — | — | — | — | |
| ATMG†Backbone=ViT-B/16, Zero-shot=true2026.03 | 91.2 | — | — | — | — | — | |
| BYOLBackbone=ResNet-50, Epochs=1000, Evaluation Protocol=linear2023.12 | 91.18 | — | — | — | — | — | |
| FastCLIPBase Model=SigLIP ViT-B/162026.01 | 91.18 | — | — | — | — | — | |
| Mixed Barlow TwinsBackbone=ResNet-50, Epochs=1000, Evaluation Protocol=linear2023.12 | 91.1 | — | — | — | — | — | |
| Text2ConceptEvaluation Protocol=zero-shot2024.11 | 91 | — | — | — | — | — | |
| ATMG†Backbone=ViT-S/16, Zero-shot=true2026.03 | 90.7 | — | — | — | — | — | |
| CLIPBackbone=ViT-S/16, Zero-shot=true2026.03 | 89.7 | — | — | — | — | — | |
| MERUBackbone=ViT-S/16, Zero-shot=true2026.03 | 89.7 | — | — | — | — | — | |
| LaCLIPArchitecture=ViT-B/16, Pre-training Data=CC3M, Protocol=5-way 5-shot2023.05 | 89.5 | — | — | — | — | — | |
| SimCLRBackbone=ResNet-50, Epochs=1000, Evaluation Protocol=linear2023.12 | 89.26 | — | — | — | — | — | |
| IIC plus finetuneSupervision=Semi-supervised2018.07 | 88.8 | — | — | — | — | — | |
| Barlow TwinsBackbone=ResNet-50, Epochs=1000, Evaluation Protocol=linear2023.12 | 87.93 | — | — | — | — | — | |
| Mixed Barlow TwinsBackbone=ResNet-50, Epochs=2000, Evaluation Protocol=k-NN2023.12 | 87.79 | — | — | — | — | — | |
| OyallonSupervision=Fully supervised2018.07 | 87.6 | — | — | — | — | — | |
| Mixed Barlow TwinsBackbone=ResNet-50, Epochs=1000, Evaluation Protocol=k-NN2023.12 | 87.55 | — | — | — | — | — | |
| CutoutSupervision=Fully supervised2018.07 | 87.3 | — | — | — | — | — | |
| SD Featureszero-shot=false2023.03 | 87.2 | — | — | — | — | — | |
| FastBUSNlabels=8002026.02 | 87.08 | — | — | — | — | — | |
| BYOLBackbone=ResNet-50, Epochs=1000, Evaluation Protocol=k-NN2023.12 | 86.94 | — | — | — | — | — | |
| BYOLAugSelf=true2021.11 | 86.79 | — | — | — | — | — | |
| BYOLAugSelf=false2021.11 | 86.73 | — | — | — | — | — | |
| Hamiltonian2017.09 | 85.5 | — | — | — | — | — | |
| SimCLRBackbone=ResNet-50, Epochs=1000, Evaluation Protocol=k-NN2023.12 | 85.14 | — | — | — | — | — | |
| SimCLRAugSelf=true2021.11 | 84.99 | — | — | — | — | — | |
| SimCLRAugSelf=false2021.11 | 84.87 | — | — | — | — | — | |
| Barlow TwinsBackbone=ResNet-50, Epochs=1000, Evaluation Protocol=k-NN2023.12 | 84.78 | — | — | — | — | — | |
| MidPoint2017.09 | 84.6 | — | — | — | — | — | |
| CLIPArchitecture=ViT-B/16, Pre-training Data=CC3M, Protocol=5-way 5-shot2023.05 | 84.6 | — | — | — | — | — | |
| Leapfrog2017.09 | 83.7 | — | — | — | — | — | |
| SwAVAugSelf=true2021.11 | 82.57 | — | — | — | — | — | |
| SwAVAugSelf=false2021.11 | 82.21 | — | — | — | — | — | |
| FastBUSNlabels=2502026.02 | 81.78 | — | — | — | — | — | |
| IIC plus finetuneEvaluation protocol=Multi-fold evaluation, Supervision=Semi-supervised2018.07 | 79.2 | — | — | — | — | — | |
| RegMeanBackbone=ViT-B/16, Pre-trained=ImageNet-21k, Number of Merged Models=302024.05 | 78.94 | — | — | — | — | — | |
| MixMatchNlabels=8002026.02 | 78.83 | — | — | — | — | — | |
| BenignPre-train Dataset=CIFAR-102026.02 | 77.23 | — | — | — | — | — | |
| DeepINFOMAX 20182018.07 | 77 | — | — | — | — | — | |
| HPEPre-train Dataset=CIFAR-102026.02 | 76.36 | — | — | — | 99.41 | — | |
| OyallonSupervision=Fully supervised, Evaluation protocol=Multi-fold evaluation2018.07 | 76 | — | — | — | — | — | |
| CTRLPre-train Dataset=CIFAR-102026.02 | 75.93 | — | — | — | 61.51 | — | |
| BADFSSPre-train Dataset=CIFAR-102026.02 | 75.51 | — | — | — | 72.41 | — | |
| BadEncoderPre-train Dataset=CIFAR-102026.02 | 75.29 | — | — | — | 64.31 | — | |
| Zhao et al. 20162017.09 | 74.3 | — | — | — | — | — | |
| SWWAE 2015Evaluation protocol=Multi-fold evaluation2018.07 | 74.3 | — | — | — | — | — | |
| Dosovitskiy 2015Evaluation protocol=Multi-fold evaluation2018.07 | 74.2 | — | — | — | — | — | |
| Dundar et al. 20152017.09 | 74.1 | — | — | — | — | — | |
| Dundar 20152018.07 | 74.1 | — | — | — | — | — |