Image Classification on Caltech-101 (Accuracy)
98.9AccuracyStableRep
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| StableRepPre-training dataset=RedCaps, Pre-training data source=Syn, Evaluation protocol=Few-shot2023.06 | 98.9 | — | |
| StableRepPre-training Data Source=Synthetic, Evaluation Protocol=5-way, 5-shot2023.06 | 98.8 | — | |
| StableRepPre-training dataset=CC12M, Pre-training data source=Syn, Evaluation protocol=Few-shot2023.06 | 98.8 | — | |
| CLIPPre-training Data Source=Real, Evaluation Protocol=5-way, 5-shot2023.06 | 98.2 | — | |
| CLIPPre-training dataset=CC12M, Pre-training data source=Real, Evaluation protocol=Few-shot2023.06 | 98.2 | — | |
| CLIPPre-training dataset=RedCaps, Pre-training data source=Real, Evaluation protocol=Few-shot2023.06 | 97.8 | — | |
| AWTShots=162024.07 | 97.12 | — | |
| StatABackbone=ViT-L/14, Keff=Very Low2025.01 | 97 | — | |
| CLIPPre-training dataset=RedCaps, Pre-training data source=Syn, Evaluation protocol=Few-shot2023.06 | 96.9 | — | |
| CLIPPre-training Data Source=Synthetic, Evaluation Protocol=5-way, 5-shot2023.06 | 96.8 | — | |
| CLIPPre-training dataset=CC12M, Pre-training data source=Syn, Evaluation protocol=Few-shot2023.06 | 96.8 | — | |
| PyramidCLIPImage Encoder=ViT-B/16, Pretrain Dataset Size=143M, Evaluation Protocol=Linear probe2022.04 | 96.5 | — | |
| PyramidCLIPImage Encoder=ViT-B/32, Pretrain Dataset Size=143M, Evaluation Protocol=Linear probe2022.04 | 96.4 | — | |
| ISyNet-N32021.09 | 96.26 | — | |
| ResNet-50+2021.09 | 96.26 | — | |
| SaliencyDecorBackbone=ResNet-182026.04 | 96.2 | — | |
| StatABackbone=ViT-L/14, Keff=Low2025.01 | 96.1 | — | |
| ISyNet-N1-S32021.09 | 96.03 | — | |
| ISyNet-N22021.09 | 95.87 | — | |
| FR-ResNet18#Samples=60002022.10 | 95.86 | — | |
| ISyNet-N1-S12021.09 | 95.74 | — | |
| ISyNet-N1-S22021.09 | 95.74 | — | |
| StatABackbone=ViT-L/14, K_eff=Medium2025.01 | 95.6 | — | |
| StatABackbone=ViT-L/14, Keff=Medium2025.01 | 95.6 | — | |
| ISyNet-N12021.09 | 95.44 | — | |
| ResNet18#Samples=60002022.10 | 95.42 | — | |
| FR-ResNet18#Samples=50002022.10 | 95.21 | — | |
| CLIPBackbone=ViT-L/142025.01 | 95.2 | — | |
| CLIPEncoder=ViT-L/14, K_eff=All2025.01 | 95.2 | — | |
| AWTShots=12024.07 | 95.13 | — | |
| ISyNet-N02021.09 | 95.05 | — | |
| ResNet-34+2021.09 | 95.05 | — | |
| StatAEncoder=ViT-L/14, K_eff=All2025.01 | 95 | — | |
| ResNet18#Samples=50002022.10 | 94.94 | — | |
| ResNet-18+2021.09 | 94.92 | — | |
| CLIP*Image Encoder=ViT-B/16, Pretrain Dataset Size=400M, Evaluation Protocol=Linear probe2022.04 | 94.7 | — | |
| SGTBackbone=ResNet-182026.04 | 94.7 | — | |
| Supervised-INArchitecture=ResNet-50, Pre-training Dataset=ImageNet, Evaluation Protocol=Linear evaluation2020.06 | 94.5 | — | |
| SUPEvaluation Protocol=Linear evaluation, Backbone=ResNet-50, Pre-trained=ImageNet2021.10 | 94.5 | — | |
| BaselineBackbone=ResNet-182026.04 | 94.5 | — | |
| FR-ResNet18#Samples=40002022.10 | 94.49 | — | |
| CoCoOpSource Dataset=ImageNet, Shots per class=16, Transfer Setting=Cross-dataset transfer2022.03 | 94.43 | — | |
| ResNet18#Samples=40002022.10 | 94.31 | — | |
| SupervisedProtocol=Fine-tuned, Backbone=ResNet-50 (4x), Pre-training=ImageNet2020.02 | 94.2 | — | |
| BYOLArchitecture=ResNet-50, Pre-training Dataset=ImageNet, Evaluation Protocol=Linear evaluation2020.06 | 94.2 | — | |
| BYOLEvaluation Protocol=Linear evaluation, Backbone=ResNet-50, Pre-trained=ImageNet2021.10 | 94.2 | — | |
| BYOLEvaluation Protocol=Linear eval, Backbone=ResNet, Pre-training Dataset=ImageNet2020.11 | 94.2 | — | |
| NLEEPTop-k=k=32022.07 | 94.12 | — | |
| SupervisedProtocol=Linear evaluation, Backbone=ResNet-50 (4x), Pre-training=ImageNet2020.02 | 94.1 | — | |
| SimCLRProtocol=Fine-tuned, Backbone=ResNet-50 (4x), Pre-training=ImageNet2020.02 | 94.1 | — | |
| OpenCLIPPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Protocol=Zero-shot classification2024.07 | 94 | — | |
| MTABatch Size=1000, Backbone=ViT-B/162025.01 | 94 | — | |
| SFDATop-k=k=32022.07 | 93.95 | — | |
| SimCLRProtocol=Linear evaluation, Backbone=ResNet-50 (4x), Pre-training=ImageNet2020.02 | 93.9 | — | |
| MLCDPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Protocol=Zero-shot classification2024.07 | 93.9 | — | |
| LogMETop-k=k=22022.07 | 93.87 | — | |
| SFDAcomTop-k=k=22022.07 | 93.87 | — | |
| LogMETop-k=k=32022.07 | 93.87 | — | |
| SFDAcomTop-k=k=32022.07 | 93.87 | — | |
| BYOLArchitecture=ResNet-50, Pre-training Dataset=ImageNet, Evaluation Protocol=Fine-tuned2020.06 | 93.8 | — | |
| BYOLEvaluation Protocol=Fine-tune, Backbone=ResNet-50, Pre-trained=ImageNet2021.10 | 93.8 | — | |
| BYOLEvaluation Protocol=Fine-tuned, Backbone=ResNet, Pre-training Dataset=ImageNet2020.11 | 93.8 | — | |
| FLIPPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Protocol=Zero-shot classification2024.07 | 93.8 | — | |
| StatABatch Size=1000, Backbone=ViT-B/16, Keff (Effective classes per task)=Very High (50-100)2025.01 | 93.8 | — | |
| CoOpSource Dataset=ImageNet, Shots per class=16, Transfer Setting=Cross-dataset transfer2022.03 | 93.7 | — | |
| TransCLIPBatch Size=1000, Backbone=ViT-B/16, Keff (Effective classes per task)=Very High (50-100)2025.01 | 93.7 | — | |
| SL-CE-tuningBackbone=ResNet-50, Pre-training=Supervised ImageNet2021.02 | 93.65 | — | |
| FNCEvaluation Protocol=Fine-tuned, Backbone=ResNet, Pre-training Dataset=ImageNet2020.11 | 93.6 | — | |
| FR-ResNet18#Samples=30002022.10 | 93.56 | — | |
| TWISTEvaluation Protocol=Fine-tune, Backbone=ResNet-50, Pre-trained=ImageNet2021.10 | 93.5 | — | |
| FNCEvaluation Protocol=Linear eval, Backbone=ResNet, Pre-training Dataset=ImageNet2020.11 | 93.5 | — | |
| Core-tuningBackbone=ResNet-50, Pre-training=MoCo-v22021.02 | 93.46 | — | |
| Core-tuningBackbone=ResNet-50, Pre-training=MoCo-v22021.02 | 93.46 | — | |
| NLEEPTop-k=k=22022.07 | 93.45 | — | |
| SFDATop-k=k=22022.07 | 93.4 | — | |
| StatABatch Size=1000, Backbone=ViT-B/16, Keff (Effective classes per task)=High (25-50)2025.01 | 93.4 | — | |
| StatABackbone=ViT-B/32, Keff=Very Low2025.01 | 93.4 | — | |
| SL-CE-tuningBackbone=ResNet-50, Pre-training=Supervised2021.02 | 93.39 | — | |
| Handcrafted2024.07 | 93.35 | — | |
| Supervised-INArchitecture=ResNet-50, Pre-training Dataset=ImageNet, Evaluation Protocol=Fine-tuned2020.06 | 93.3 | — | |
| SUPEvaluation Protocol=Fine-tune, Backbone=ResNet-50, Pre-trained=ImageNet2021.10 | 93.3 | — | |
| CLIPBatch Size=1000, Backbone=ViT-B/162025.01 | 93.2 | — | |
| ZLaPBatch Size=1000, Backbone=ViT-B/16, Keff (Effective classes per task)=High (25-50)2025.01 | 93.2 | — | |
| StatABackbone=ViT-B/32, Keff=Low2025.01 | 93.2 | — | |
| CLIP*Image Encoder=ViT-B/32, Pretrain Dataset Size=400M, Evaluation Protocol=Linear probe2022.04 | 93 | — | |
| EsViTBackbone=Swin-T, Evaluation Protocol=Linear probe2021.06 | 93 | — | |
| M&MBackbone=ResNet-50, Pre-training=MoCo-v22021.02 | 92.91 | — | |
| M&MBackbone=ResNet-50, Pre-training=MoCo-v22021.02 | 92.91 | — | |
| StatABackbone=ViT-B/32, Keff=Medium2025.01 | 92.9 | — | |
| SCLBackbone=ResNet-50, Pre-training=MoCo-v22021.02 | 92.84 | — | |
| SCLBackbone=ResNet-50, Pre-training=MoCo-v22021.02 | 92.84 | — | |
| Bi-tuningBackbone=ResNet-50, Pre-training=MoCo-v22021.02 | 92.83 | — | |
| Bi-tuningBackbone=ResNet-50, Pre-training=MoCo-v22021.02 | 92.83 | — | |
| TransCLIPBatch Size=1000, Backbone=ViT-B/16, Keff (Effective classes per task)=Medium (5-25)2025.01 | 92.8 | — | |
| StatABackbone=ResNet-101, Keff=Low2025.01 | 92.8 | — | |
| ResNet18#Samples=30002022.10 | 92.78 | — | |
| CLIP (reported)Pre-training Data=WIT-400M, Backbone=ViT-L/14, Evaluation Protocol=Zero-shot classification2024.07 | 92.6 | — | |
| CLIPPre-training Data=WIT-400M, Backbone=ViT-L/14, Evaluation Protocol=Zero-shot classification2024.07 | 92.4 | — | |
| SimCLR v2Evaluation Protocol=Fine-tuned, Backbone=ResNet, Pre-training Dataset=ImageNet2020.11 | 92.3 | — | |
| DELTABackbone=ResNet-50, Pre-training=MoCo-v22021.02 | 92.19 | — |