Image Classification on Oxford-IIIT Pets (Top-1 Accuracy)
96.4Top-1 AccFlorence-CoSwin-H
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Florence-CoSwin-HResolution=384px, Backbone=CoSwin-H, Evaluation Protocol=Linear Probing2021.11 | 96.4 | — | |
| SynCLR + GMAILEvaluation Protocol=Zero-shot2026.02 | 95.7 | — | |
| EfficientNet-L2Resolution=800px, Evaluation Protocol=Linear Probing2021.11 | 95.6 | — | |
| LOUPEevaluation=linear probing2022.08 | 95.5 | — | |
| ProDA2023.11 | 95.43 | — | |
| MaPLe2023.11 | 95.43 | — | |
| PromptAlign2023.11 | 95.38 | — | |
| MaPLe+TPTtest-time prompt tuning=true2023.11 | 95.23 | — | |
| CLIP + GMAILEvaluation Protocol=Zero-shot2026.02 | 95.23 | — | |
| CLIP-ViT-L/14Resolution=336px, Evaluation Protocol=Linear Probing2021.11 | 95.1 | — | |
| CLIPevaluation=linear probing2022.08 | 95.1 | — | |
| CLIP-ResNet-50x64Evaluation Protocol=Linear Probing2021.11 | 94.5 | — | |
| LOUPEmode=zero-shot2022.08 | 94.1 | — | |
| SynCLREvaluation Protocol=Zero-shot2026.02 | 93.6 | — | |
| CLIPmode=zero-shot2022.08 | 93.5 | — | |
| CLIPEvaluation Protocol=Zero-shot2026.02 | 93.33 | — | |
| LoRA (r=1)Backbone=ViT-B, Byte Footprint=74KB2026.04 | 93 | — | |
| ViT-L/16Resolution=384px, Evaluation Protocol=Linear Probing2021.11 | 92.9 | — | |
| SimCLRv2Backbone=ResNet-152x3, Evaluation Protocol=Linear Probing2021.11 | 92.6 | — | |
| SOLAR r=1(2→0.2)Backbone=ViT-B, Byte Footprint=8KB (89% ↓)2026.04 | 92.6 | — | |
| AWTTrain=false2024.07 | 92.53 | — | |
| LlipPre-training Data=2.5B, Model Architecture=ViT-B/32, Zero-shot=true2024.12 | 92.3 | — | |
| BYOLEvaluation Protocol=Fine-tuned, Backbone=ResNet, Pre-training Dataset=ImageNet2020.11 | 91.7 | — | |
| ProVP-RefTrain=true2024.07 | 91.58 | — | |
| Self-TPT-vTrain=true2024.07 | 91.26 | — | |
| FNCEvaluation Protocol=Fine-tuned, Backbone=ResNet, Pre-training Dataset=ImageNet2020.11 | 90.9 | — | |
| MetaCLIPPre-training Data=2.5B, Model Architecture=ViT-B/32, Zero-shot=true2024.12 | 90.9 | — | |
| PromptAlignTrain=true2024.07 | 90.76 | — | |
| SuS-X-SDTrain=false2024.07 | 90.57 | — | |
| OpenCLIPImage Encoder=ViT-B/16, Zero-shot=true2026.02 | 90.5 | — | |
| MaPLeTrain=true2024.07 | 90.49 | — | |
| PLOT++Train=true2024.07 | 90.49 | — | |
| BYOLEvaluation Protocol=Linear eval, Backbone=ResNet, Pre-training Dataset=ImageNet2020.11 | 90.4 | — | |
| DST (FlexMatch)Pre-training=Supervised, Backbone=ResNet-50, Labels per category=42022.02 | 90.4 | — | |
| OpenCLIPPre-training Data=LAION-2B, Model Architecture=ViT-B/32, Zero-shot=true2024.12 | 90.4 | — | |
| NOLABackbone=ViT-B, Byte Footprint=48KB2026.04 | 90.4 | — | |
| OpenCLIPPre-training Data=DataComp-1B, Model Architecture=ViT-B/32, Zero-shot=true2024.12 | 90.2 | — | |
| CoCoOpTrain=true2024.07 | 90.14 | — | |
| WaffleCLIPTrain=false2024.07 | 89.95 | — | |
| SimCLR v2Evaluation Protocol=Fine-tuned, Backbone=ResNet, Pre-training Dataset=ImageNet2020.11 | 89.9 | — | |
| DST (FixMatch)Pre-training=Supervised, Backbone=ResNet-50, Labels per category=42022.02 | 89.8 | — | |
| SimCLR v1Evaluation Protocol=Fine-tuned, Backbone=ResNet, Pre-training Dataset=ImageNet2020.11 | 89.2 | — | |
| CoOpTrain=true2024.07 | 89.14 | — | |
| CuPLTrain=false2024.07 | 89.13 | — | |
| CLIP-ViT-B/16Layers=-2026.03 | 89.13 | — | |
| POMPTrain=true2024.07 | 89.05 | — | |
| FNCEvaluation Protocol=Linear eval, Backbone=ResNet, Pre-training Dataset=ImageNet2020.11 | 89 | — | |
| VisDescTrain=false2024.07 | 88.85 | — | |
| CLIPTrain=false2024.07 | 88.25 | — | |
| DiffTPTTrain=true2024.07 | 88.22 | — | |
| TPTTrain=true2024.07 | 87.79 | — | |
| Col-LnLayers=3, 6, 9, KeepRate=0.72026.03 | 85.99 | — | |
| CLIPImage Encoder=ViT-B/16, Zero-shot=true2026.02 | 85.2 | — | |
| TCA (Baseline)Layers=3, 6, 9, KeepRate=0.72026.03 | 85.15 | — | |
| SimCLR v2Evaluation Protocol=Linear eval, Backbone=ResNet, Pre-training Dataset=ImageNet2020.11 | 85 | — | |
| OpenCLIPPre-training Data=LAION-400M, Model Architecture=ViT-B/32, Zero-shot=true2024.12 | 84.9 | — | |
| SimCLR v1Evaluation Protocol=Linear eval, Backbone=ResNet, Pre-training Dataset=ImageNet2020.11 | 83.6 | — | |
| Noun SubmanifoldBackbone=CLIP ViT-L-14, zero-shot=true2023.05 | 83.1 | — | |
| CLIP-MapbaseImage Encoder=ViT-39M/16, Zero-shot=true2026.02 | 83.1 | — | |
| POS PGABackbone=CLIP ViT-L-14, zero-shot=true2023.05 | 82.6 | — | |
| CLIPBackbone=CLIP ViT-L-14, zero-shot=true2023.05 | 82.5 | — | |
| POS PCABackbone=CLIP ViT-L-14, zero-shot=true2023.05 | 82.5 | — | |
| Col-LnLayers=0, 3, 6, KeepRate=0.72026.03 | 82.47 | — | |
| TCA (Baseline)Layers=0, 3, 6, KeepRate=0.72026.03 | 82.12 | — | |
| TinyCLIPImage Encoder=ViT-39M/16, Zero-shot=true2026.02 | 80.8 | — | |
| BaselinePre-training=Supervised, Backbone=ResNet-50, Labels per category=42022.02 | 78.1 | — | |
| DST (FixMatch)Pre-training=Unsupervised, Backbone=ResNet-50, Labels per category=42022.02 | 75.4 | — | |
| DST (FlexMatch)Pre-training=Unsupervised, Backbone=ResNet-50, Labels per category=42022.02 | 75.1 | — | |
| OURS_SepProtocol=Linear probing2023.03 | 67.16 | — | |
| OURS_BrProtocol=Linear probing2023.03 | 65.71 | — | |
| OURS_GCProtocol=Linear probing2023.03 | 64.51 | — | |
| DreamLIPPre-training Data=Merged-30M, Model Architecture=ViT-B/32, Zero-shot=true2024.12 | 62.8 | — | |
| CyCLIPProtocol=Linear probing2023.03 | 62.63 | — | |
| BaselinePre-training=Unsupervised, Backbone=ResNet-50, Labels per category=42022.02 | 60 | — | |
| CLIPProtocol=Linear probing2023.03 | 59.66 | — | |
| COSMOSPre-training Data=Merged-30M, Model Architecture=ViT-B/32, Zero-shot=true2024.12 | 57.3 | — | |
| COSMOSPre-training Data=CC12M, Model Architecture=ViT-B/32, Zero-shot=true2024.12 | 54.2 | — | |
| CLIP-MapsmallImage Encoder=ViT-8M/16, Zero-shot=true2026.02 | 50.9 | — | |
| CLIPPre-training Data=Merged-30M, Model Architecture=ViT-B/32, Zero-shot=true2024.12 | 49.4 | — | |
| SigLIPPre-training Data=Merged-30M, Model Architecture=ViT-B/32, Zero-shot=true2024.12 | 47.8 | — | |
| †TinyCLIPImage Encoder=ViT-8M/16, Zero-shot=true2026.02 | 46.2 | — | |
| CLIPPre-training Data=CC12M, Model Architecture=ViT-B/32, Zero-shot=true2024.12 | 43.3 | — | |
| DreamLIPPre-training Data=CC12M, Model Architecture=ViT-B/32, Zero-shot=true2024.12 | 41.9 | — | |
| SigLIPPre-training Data=CC12M, Model Architecture=ViT-B/32, Zero-shot=true2024.12 | 41.8 | — | |
| COSMOSPre-training Data=YFCC15M, Model Architecture=ViT-B/32, Zero-shot=true2024.12 | 38.6 | — | |
| DreamLIPPre-training Data=YFCC15M, Model Architecture=ViT-B/32, Zero-shot=true2024.12 | 36.3 | — | |
| SigLIPPre-training Data=YFCC15M, Model Architecture=ViT-B/32, Zero-shot=true2024.12 | 30.4 | — | |
| COSMOSPre-training Data=CC3M, Model Architecture=ViT-B/32, Zero-shot=true2024.12 | 29.9 | — | |
| CLIPPre-training Data=YFCC15M, Model Architecture=ViT-B/32, Zero-shot=true2024.12 | 27.6 | — | |
| DreamLIPPre-training Data=CC3M, Model Architecture=ViT-B/32, Zero-shot=true2024.12 | 18.8 | — | |
| CLIP-MaptinyImage Encoder=ViT-0.8M/16, Zero-shot=true2026.02 | 18.8 | — | |
| †TinyCLIPImage Encoder=ViT-0.8M/16, Zero-shot=true2026.02 | 17.4 | — | |
| CLIPPre-training Data=CC3M, Model Architecture=ViT-B/32, Zero-shot=true2024.12 | 16.9 | — | |
| SigLIPPre-training Data=CC3M, Model Architecture=ViT-B/32, Zero-shot=true2024.12 | 16.3 | — | |
| CEConv-32024.06 | — | 31.08 | |
| CEConv-42024.06 | — | 33.7 | |
| Hue-3*2024.06 | — | 31.37 | |
| Hue-4-Sat-3*2024.06 | — | 29.84 | |
| Hue-4*2024.06 | — | 27.39 | |
| ResNet2024.06 | — | 31.52 |