Image Classification on Caltech-101
98.6Top-1 AccuracyPromptAlign
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| PromptAlign2023.11 | 98.6 | — | — | |
| MaPLe+TPTtest-time prompt tuning=true2023.11 | 98.54 | — | — | |
| ProDA2023.11 | 98.27 | — | — | |
| MaPLe2023.11 | 98 | — | — | |
| LOUPEevaluation=linear probing2022.08 | 97.5 | — | — | |
| Florence-CoSwin-HResolution=384px, Backbone=CoSwin-H, Evaluation Protocol=Linear Probing2021.11 | 96.6 | — | — | |
| CLIPevaluation=linear probing2022.08 | 96.5 | — | — | |
| PEFTShots=102023.05 | 96.1 | — | — | |
| CLIP-ViT-L/14Resolution=336px, Evaluation Protocol=Linear Probing2021.11 | 96 | — | — | |
| OURSShots=8, Backbone=ViT-B/16, Templates=72025.12 | 96 | — | — | |
| OURSShots=8, Backbone=ViT-B/16, Templates=12025.12 | 95.9 | — | — | |
| CLIP-LoRAShots=8, Backbone=ViT-B/16, Templates=72025.12 | 95.8 | — | — | |
| OURSShots=4, Backbone=ViT-B/16, Templates=12025.12 | 95.7 | — | — | |
| CLIP-LoRAShots=8, Backbone=ViT-B/16, Templates=12025.12 | 95.6 | — | — | |
| OURSShots=4, Backbone=ViT-B/16, Templates=72025.12 | 95.5 | — | — | |
| CLIP-ResNet-50x64Evaluation Protocol=Linear Probing2021.11 | 95.4 | — | — | |
| LoRA-CLIPShots=4, Backbone=ViT-B/16, Templates=12025.12 | 95.2 | — | — | |
| LoRA-CLIPShots=4, Backbone=ViT-B/16, Templates=72025.12 | 95.1 | — | — | |
| SimCLRv2Backbone=ResNet-152x3, Evaluation Protocol=Linear Probing2021.11 | 94.9 | — | — | |
| OURSShots=1, Backbone=ViT-B/16, Templates=12025.12 | 94.8 | — | — | |
| ViT-L/16Resolution=384px, Evaluation Protocol=Linear Probing2021.11 | 94.7 | — | — | |
| EfficientNet-L2Resolution=800px, Evaluation Protocol=Linear Probing2021.11 | 94.7 | — | — | |
| LlipData Size=2.5B, Vision Encoder=ViT-B/16, Evaluation Protocol=Zero-shot2024.12 | 94.7 | — | — | |
| CoOpShots=102023.05 | 94.6 | — | — | |
| PEFTShots=52023.05 | 94.5 | — | — | |
| OURSShots=1, Backbone=ViT-B/16, Templates=72025.12 | 94.5 | — | — | |
| CaCoBackbone=ResNet-50, Pre-training Dataset=ImageNet-1K, Pre-training Epochs=8002022.03 | 94.4 | — | — | |
| BYOLBackbone=ResNet-50, Pre-training Dataset=ImageNet-1K, Pre-training Epochs=8002022.03 | 94.2 | — | — | |
| LoRA-CLIPShots=1, Backbone=ViT-B/16, Templates=72025.12 | 94.1 | — | — | |
| LOUPEmode=zero-shot2022.08 | 93.9 | — | — | |
| LoRA-CLIPShots=1, Backbone=ViT-B/16, Templates=12025.12 | 93.7 | — | — | |
| LaFTerBackbone=ViT-B/322023.05 | 93.3 | — | — | |
| LaFTerShots=02023.05 | 93.3 | — | — | |
| MetaCLIPData Size=2.5B, Vision Encoder=ViT-B/16, Evaluation Protocol=Zero-shot2024.12 | 93.3 | — | — | |
| CoOpShots=52023.05 | 93.2 | — | — | |
| OpenCLIPData Size=2B, Vision Encoder=ViT-B/16, Evaluation Protocol=Zero-shot2024.12 | 93.2 | — | — | |
| LlipPre-training Data=2.5B, Model Architecture=ViT-B/32, Zero-shot=true2024.12 | 92.9 | — | — | |
| MetaCLIPPre-training Data=2.5B, Model Architecture=ViT-B/32, Zero-shot=true2024.12 | 92.8 | — | — | |
| CLIPmode=zero-shot2022.08 | 92.6 | — | — | |
| WukongBackbone=ViT-L, Zero-shot=true, Pre-training Dataset=Wukong2022.02 | 92.4 | — | — | |
| CLIPBackbone=ViT-L, Zero-shot=true, Pre-training Dataset=Wukong2022.02 | 91.9 | — | — | |
| AdCo +Backbone=ResNet-50, Pre-training Dataset=ImageNet-1K, Pre-training Epochs=8002022.03 | 91.8 | — | — | |
| CoOpShots=12023.05 | 91.7 | — | — | |
| FILIPBackbone=Swin-L, Zero-shot=true, Pre-training Dataset=Wukong2022.02 | 91.6 | — | — | |
| xCLIPLinear probing=true, Backbone=ViT-B/16, Pre-train=IT35M2022.10 | 91.6 | — | — | |
| OpenCLIPPre-training Data=DataComp-1B, Model Architecture=ViT-B/32, Zero-shot=true2024.12 | 91.6 | — | — | |
| MoCo v2+Backbone=ResNet-50, Pre-training Dataset=ImageNet-1K, Pre-training Epochs=8002022.03 | 91.4 | — | — | |
| NNCLRBackbone=ResNet-50, Pre-training Dataset=ImageNet-1K, Pre-training Epochs=8002022.03 | 91.3 | — | — | |
| DebiasMatchPre-training=Supervised, Backbone=ResNet-50, Labels per category=42022.02 | 91 | — | — | |
| CLIPBackbone=Swin-L, Zero-shot=true, Pre-training Dataset=Wukong2022.02 | 90.7 | — | — | |
| DST (FlexMatch)Pre-training=Supervised, Backbone=ResNet-50, Labels per category=42022.02 | 90.6 | — | — | |
| UPLBackbone=ViT-B/322023.05 | 90.6 | — | — | |
| PEFTShots=12023.05 | 90.6 | — | — | |
| nCLIPLinear probing=true, Backbone=ViT-B/16, Pre-train=IT35M2022.10 | 90.5 | — | — | |
| CLIPBackbone=ViT-B/322023.05 | 90.5 | — | — | |
| OpenCLIPPre-training Data=LAION-2B, Model Architecture=ViT-B/32, Zero-shot=true2024.12 | 90.5 | — | — | |
| DST (FlexMatch)Pre-training=Unsupervised, Backbone=ResNet-50, Labels per category=42022.02 | 90.4 | — | — | |
| DST (FixMatch)Pre-training=Unsupervised, Backbone=ResNet-50, Labels per category=42022.02 | 90.1 | — | — | |
| CLIPLinear probing=true, Backbone=ViT-B/16, Pre-train=IT35M2022.10 | 90 | — | — | |
| FILIPBackbone=ViT-L, Zero-shot=true, Pre-training Dataset=Wukong2022.02 | 89.9 | — | — | |
| WukongBackbone=Swin-L, Zero-shot=true, Pre-training Dataset=Wukong2022.02 | 89.8 | — | — | |
| DST (FixMatch)Pre-training=Supervised, Backbone=ResNet-50, Labels per category=42022.02 | 89.6 | — | — | |
| SimCLRBackbone=ResNet-50, Pre-training Dataset=ImageNet-1K, Pre-training Epochs=8002022.03 | 89.3 | — | — | |
| CLIPMode=Zero-shot2024.04 | 89.3 | — | — | |
| CLIPBackbone=ViT-B, Zero-shot=true, Pre-training Dataset=Wukong2022.02 | 89.2 | — | — | |
| WukongBackbone=ViT-B, Zero-shot=true, Pre-training Dataset=Wukong2022.02 | 89.1 | — | — | |
| Self-TuningPre-training=Supervised, Backbone=ResNet-50, Labels per category=42022.02 | 88.6 | — | — | |
| ARFFinetuning Dataset=ImageNet2024.04 | 88.6 | — | — | |
| OpenCLIPPre-training Data=LAION-400M, Model Architecture=ViT-B/32, Zero-shot=true2024.12 | 88.4 | — | — | |
| DreamLIPData Size=30M, Vision Encoder=ViT-B/16, Evaluation Protocol=Zero-shot2024.12 | 88 | — | — | |
| FLYPFinetuning Dataset=ImageNet2024.04 | 87.6 | — | — | |
| FlexMatchPre-training=Supervised, Backbone=ResNet-50, Labels per category=42022.02 | 87.1 | — | — | |
| FLAIRData Size=30M, Vision Encoder=ViT-B/16, Evaluation Protocol=Zero-shot2024.12 | 86.5 | — | — | |
| FlexMatchPre-training=Unsupervised, Backbone=ResNet-50, Labels per category=42022.02 | 86.4 | — | — | |
| DebiasMatchPre-training=Unsupervised, Backbone=ResNet-50, Labels per category=42022.02 | 86.4 | — | — | |
| Pseudo LabelPre-training=Supervised, Backbone=ResNet-50, Labels per category=42022.02 | 86.3 | — | — | |
| FixMatchPre-training=Supervised, Backbone=ResNet-50, Labels per category=42022.02 | 86.3 | — | — | |
| Pseudo LabelPre-training=Unsupervised, Backbone=ResNet-50, Labels per category=42022.02 | 86.2 | — | — | |
| DreamLIPPre-training Data=Merged-30M, Model Architecture=ViT-B/32, Zero-shot=true2024.12 | 86.1 | — | — | |
| COSMOSPre-training Data=Merged-30M, Model Architecture=ViT-B/32, Zero-shot=true2024.12 | 86.1 | — | — | |
| UDAPre-training=Supervised, Backbone=ResNet-50, Labels per category=42022.02 | 85.8 | — | — | |
| COSMOSPre-training Data=CC12M, Model Architecture=ViT-B/32, Zero-shot=true2024.12 | 85.5 | — | — | |
| MixMatchPre-training=Supervised, Backbone=ResNet-50, Labels per category=42022.02 | 85.4 | — | — | |
| UDAPre-training=Unsupervised, Backbone=ResNet-50, Labels per category=42022.02 | 85 | — | — | |
| CLIP-PRBackbone=ViT-B/322023.05 | 84.8 | — | — | |
| OURS_SepProtocol=Linear probing2023.03 | 84.45 | — | — | |
| SigLIPPre-training Data=Merged-30M, Model Architecture=ViT-B/32, Zero-shot=true2024.12 | 84.4 | — | — | |
| SigLIPData Size=30M, Vision Encoder=ViT-B/16, Evaluation Protocol=Zero-shot2024.12 | 84.3 | — | — | |
| VATPre-training=Supervised, Backbone=ResNet-50, Labels per category=42022.02 | 84.1 | — | — | |
| MixMatchPre-training=Unsupervised, Backbone=ResNet-50, Labels per category=42022.02 | 84.1 | — | — | |
| RATPre-training=Supervised, Backbone=ResNet-50, Labels per category=42022.02 | 84 | — | — | |
| LP-FTFinetuning Dataset=ImageNet2024.04 | 84 | — | — | |
| CLIPPre-training Data=Merged-30M, Model Architecture=ViT-B/32, Zero-shot=true2024.12 | 83.8 | — | — | |
| Mean TeacherPre-training=Supervised, Backbone=ResNet-50, Labels per category=42022.02 | 83.7 | — | — | |
| CLIPData Size=30M, Vision Encoder=ViT-B/16, Evaluation Protocol=Zero-shot2024.12 | 83.7 | — | — | |
| PI-ModelPre-training=Supervised, Backbone=ResNet-50, Labels per category=42022.02 | 83.5 | — | — | |
| OURS_GCProtocol=Linear probing2023.03 | 83.23 | — | — | |
| DreamLIPPre-training Data=CC12M, Model Architecture=ViT-B/32, Zero-shot=true2024.12 | 83.2 | — | — | |
| FILIPBackbone=ViT-B, Zero-shot=true, Pre-training Dataset=Wukong2022.02 | 83.1 | — | — | |
| FixMatchPre-training=Unsupervised, Backbone=ResNet-50, Labels per category=42022.02 | 83.1 | — | — |