Image Classification on Oxford-IIIT Pet (Top-1 accuracy)
95.9Top-1 AccuracyPiSSA
Evaluation Results
| Method | Links | |
|---|---|---|
| PiSSABackbone=ViT-Base, Budget=100%2026.05 | 95.9 | |
| PiSSABackbone=ViT-Large, Budget=100%2026.05 | 95.7 | |
| SigLIP2Data size=10B, Zero-shot=true2026.03 | 95.4 | |
| JACTUSBackbone=ViT-Large, Budget=80%2026.05 | 95.2 | |
| LoRABackbone=ViT-Large, Budget=100%2026.05 | 94.8 | |
| DoRABackbone=ViT-Large, Budget=100%2026.05 | 94.8 | |
| PEcoreData size=5B, Zero-shot=true2026.03 | 94.6 | |
| AdapterTuneBackbone=ViT-B/16, Trainable Parameters (%)=0.9%2026.03 | 94.3 | |
| SigLIPData size=10B, Zero-shot=true2026.03 | 94.2 | |
| SVD-LLM+LoRABackbone=ViT-Large, Budget=80%2026.05 | 94.2 | |
| DoRABackbone=ViT-Base, Budget=100%2026.05 | 93.8 | |
| AdapterTuneBackbone=ViT-S/16, Trainable Parameters (%)=0.9%2026.03 | 93.5 | |
| LlipData size=2.5B, Zero-shot=true2026.03 | 93.5 | |
| JACTUSBackbone=ViT-Base, Budget=80%2026.05 | 93.4 | |
| Head-only tuningBackbone=ViT-B/16, Trainable Parameters (%)=0.1%2026.03 | 93.3 | |
| SVFTBackbone=ViT-Large, Budget=100%2026.05 | 93.3 | |
| JACTUSBackbone=ViT-Large, Budget=60%2026.05 | 93.2 | |
| LoRABackbone=ViT-Base, Budget=100%2026.05 | 93.1 | |
| SVD-LLM+LoRABackbone=ViT-Base, Budget=80%2026.05 | 92.7 | |
| CapPaEvaluation Protocol=10-shot linear evaluation, Model Scale=L/142023.06 | 92.6 | |
| SVFTBackbone=ViT-Base, Budget=100%2026.05 | 92.5 | |
| SVD-LLM+LoRABackbone=ViT-Large, Budget=60%2026.05 | 92.2 | |
| CaCoBackbone=ResNet-50, Pre-training Dataset=ImageNet-1K, Pre-training Epochs=8002022.03 | 91.9 | |
| NNCLRBackbone=ResNet-50, Pre-training Dataset=ImageNet-1K, Pre-training Epochs=8002022.03 | 91.8 | |
| IN-RealFakeBackbone=ViT-S, Pre-Trained Data=IN-RealFake, Evaluation Protocol=Linear Probing2024.12 | 91.7 | |
| MetaCLIPData size=2.5B, Zero-shot=true2026.03 | 91.7 | |
| JACTUSBackbone=ViT-Base, Budget=60%2026.05 | 91.6 | |
| AdapterTuneBackbone=DeiT-T, Trainable Parameters (%)=0.9%2026.03 | 91.4 | |
| Head-only tuningBackbone=DeiT-T, Trainable Parameters (%)=0.1%2026.03 | 90.8 | |
| IN-RealBackbone=ResNet50, Pre-Trained Data=IN-Real, Evaluation Protocol=Linear Probing2024.12 | 90.7 | |
| Head-only tuningBackbone=ViT-S/16, Trainable Parameters (%)=0.1%2026.03 | 90.5 | |
| BYOLBackbone=ResNet-50, Pre-training Dataset=ImageNet-1K, Pre-training Epochs=8002022.03 | 90.4 | |
| CLIPEvaluation Protocol=10-shot linear evaluation, Model Scale=L/142023.06 | 90.4 | |
| IN-RealBackbone=ViT-S, Pre-Trained Data=IN-Real, Evaluation Protocol=Linear Probing2024.12 | 90.2 | |
| VP (white-box)Shots=16, Model Access Type=White-box2024.07 | 90.2 | |
| SVD-LLM+LoRABackbone=ViT-Base, Budget=60%2026.05 | 90.2 | |
| IN-RealFakeBackbone=ResNet50, Pre-Trained Data=IN-RealFake, Evaluation Protocol=Linear Probing2024.12 | 90 | |
| BlackVIP-SEShots=16, Model Access Type=Black-box2024.07 | 89.8 | |
| BlackVIPShots=16, Model Access Type=Black-box2024.07 | 89.7 | |
| Full fine-tuningBackbone=ViT-S/16, Trainable Parameters (%)=100%2026.03 | 89.5 | |
| OpenCLIPData size=2B, Zero-shot=true2026.03 | 89.5 | |
| JACTUSBackbone=ViT-Large, Budget=40%2026.05 | 89.4 | |
| ZSEvaluation Protocol=Zero-shot2024.07 | 89.1 | |
| Full fine-tuningBackbone=DeiT-T, Trainable Parameters (%)=100%2026.03 | 89 | |
| lbGenBackbone=ViT-S, Pre-Trained Data=IN-lbGen, Evaluation Protocol=Linear Probing2024.12 | 88.6 | |
| BARShots=16, Model Access Type=Black-box2024.07 | 88.6 | |
| JACTUSBackbone=ViT-Base, Budget=40%2026.05 | 88.2 | |
| CLIP*Evaluation Protocol=10-shot linear evaluation, Model Scale=L/142023.06 | 87.7 | |
| MoCo v2+Backbone=ResNet-50, Pre-training Dataset=ImageNet-1K, Pre-training Epochs=8002022.03 | 87.6 | |
| lbGenBackbone=ResNet50, Pre-Trained Data=IN-lbGen, Evaluation Protocol=Linear Probing2024.12 | 87.2 | |
| VP w/ SPSA-GCShots=16, Model Access Type=Black-box2024.07 | 87.1 | |
| AdCo +Backbone=ResNet-50, Pre-training Dataset=ImageNet-1K, Pre-training Epochs=8002022.03 | 86.8 | |
| Full fine-tuningBackbone=ViT-B/16, Trainable Parameters (%)=100%2026.03 | 86.6 | |
| CapPaEvaluation Protocol=10-shot linear evaluation2023.06 | 86.5 | |
| SVD-LLM+LoRABackbone=ViT-Large, Budget=40%2026.05 | 85.8 | |
| IN-SD1.5Backbone=ViT-S, Pre-Trained Data=IN-SD1.5, Evaluation Protocol=Linear Probing2024.12 | 85.6 | |
| IN-SD1.5Backbone=ResNet50, Pre-Trained Data=IN-SD1.5, Evaluation Protocol=Linear Probing2024.12 | 85.5 | |
| SimCLRBackbone=ResNet-50, Pre-training Dataset=ImageNet-1K, Pre-training Epochs=8002022.03 | 84.6 | |
| CapEvaluation Protocol=10-shot linear evaluation2023.06 | 83.7 | |
| SVD-LLM+LoRABackbone=ViT-Base, Budget=40%2026.05 | 83.4 | |
| IN-GenRobustBackbone=ResNet50, Pre-Trained Data=IN-GenRobust, Evaluation Protocol=Linear Probing2024.12 | 83.3 | |
| IN-GenRobustBackbone=ViT-S, Pre-Trained Data=IN-GenRobust, Evaluation Protocol=Linear Probing2024.12 | 82.1 | |
| CLIPEvaluation Protocol=10-shot linear evaluation2023.06 | 82.1 | |
| CLIP*Evaluation Protocol=10-shot linear evaluation, Batch Size=16k2023.06 | 80.6 | |
| CLIP*Evaluation Protocol=10-shot linear evaluation, Batch Size=8k2023.06 | 77.7 | |
| Noun SubmanifoldBackbone=CLIP ViT-B-32, Evaluation Protocol=Zero-shot2023.05 | 75.2 | |
| POS PGABackbone=CLIP ViT-B-32, Evaluation Protocol=Zero-shot2023.05 | 75.1 | |
| CLIPBackbone=CLIP ViT-B-32, Evaluation Protocol=Zero-shot2023.05 | 73.8 | |
| POS PCABackbone=CLIP ViT-B-32, Evaluation Protocol=Zero-shot2023.05 | 73.8 | |
| DreamLIPData size=30M, Zero-shot=true2026.03 | 64.1 | |
| COSMOSData size=30M, Zero-shot=true2026.03 | 62.9 | |
| GoldiCLIPData size=30M, Zero-shot=true2026.03 | 62.7 | |
| FLAIRData size=30M, Zero-shot=true2026.03 | 55.6 | |
| SigLIPData size=30M, Zero-shot=true2026.03 | 53.3 | |
| CLIPData size=30M, Zero-shot=true2026.03 | 49.3 |