Image Classification on 11-Dataset Few-shot Learning Suite (16-shot)
86.5Average Accuracy (16-shot)SigLIP2
Evaluation Results
| Method | Links | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| SigLIP2Classifier=LDA, Embedding Projection=Original, Backbone=ViT-B2026.03 | 86.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SigLIP2Classifier=LDA, Embedding Projection=PCA<-, Backbone=ViT-B2026.03 | 86.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DroPLeModel=ViT-L/142025.12 | 86.4 | 80.4 | 98.3 | 96.3 | 89.1 | 98.9 | 93.5 | 60.8 | 82.6 | 76.7 | 84.6 | 88.9 | — | |
| SigLIPClassifier=LDA, Embedding Projection=Original, Backbone=ViT-B2026.03 | 85.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SigLIPClassifier=LDA, Embedding Projection=PCA<-, Backbone=ViT-B2026.03 | 85.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SigLIP2Classifier=Prototype, Embedding Projection=PCA<-, Backbone=ViT-B2026.03 | 84.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MMAModel=ViT-L/142025.12 | 84.4 | 79.9 | 97.6 | 95.5 | 88 | 98.4 | 92 | 56.4 | 80.2 | 75.8 | 76.3 | 88 | — | |
| CoOpModel=ViT-L/142025.12 | 84 | 78.1 | 97.5 | 94.5 | 87.4 | 98.6 | 90.2 | 53 | 77.9 | 73.7 | 86.7 | 86.7 | — | |
| SigLIPClassifier=Prototype, Embedding Projection=PCA<-, Backbone=ViT-B2026.03 | 83.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DINOv2Classifier=LDA, Embedding Projection=Original, Backbone=ViT-B2026.03 | 83.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MaPLeModel=ViT-L/142025.12 | 83.1 | 78.4 | 97.2 | 95.4 | 83.6 | 97.4 | 92 | 46.3 | 78.8 | 72.7 | 85.4 | 86.5 | — | |
| SigLIP2Classifier=Prototype, Embedding Projection=Original, Backbone=ViT-B2026.03 | 83 | — | — | — | — | — | — | — | — | — | — | — | — | |
| KgCoOpModel=ViT-L/142025.12 | 82.6 | 76.8 | 97.4 | 95.3 | 83.2 | 96.4 | 91.7 | 47.5 | 76.7 | 73.6 | 83.6 | 86.4 | — | |
| SigLIPClassifier=Prototype, Embedding Projection=Original, Backbone=ViT-B2026.03 | 82.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| CoCoOpModel=ViT-L/142025.12 | 81.7 | 77.8 | 97.4 | 95.4 | 82.7 | 95.3 | 91.9 | 45.2 | 76.7 | 71.4 | 79.8 | 85.2 | — | |
| CLIPClassifier=LDA, Embedding Projection=PCA<-, Backbone=ViT-B2026.03 | 79.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| CLIPClassifier=LDA, Embedding Projection=Original, Backbone=ViT-B2026.03 | 79.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DroPLeModel=ViT-B/322025.12 | 78.9 | 68.7 | 96.7 | 92.8 | 75.6 | 94.8 | 84.5 | 40.3 | 76.3 | 71.7 | 83.5 | 83.4 | — | |
| DINOv2Classifier=Prototype, Embedding Projection=Original, Backbone=ViT-B2026.03 | 78.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| CLIPClassifier=Prototype, Embedding Projection=PCA<-, Backbone=ViT-B2026.03 | 77.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MMAModel=ViT-B/322025.12 | 76.6 | 68 | 95.6 | 91.5 | 73.5 | 94.3 | 81.4 | 34 | 74 | 68.9 | 80.1 | 81.7 | — | |
| HOSO-AdapterBackbone=ResNet-50, Shots=162026.03 | 75.25 | 62.93 | 93.03 | 89.47 | 73.8 | 95.07 | 80.93 | 34.6 | 69.83 | 66.77 | — | 78.03 | 83.27 | |
| CoOpModel=ViT-B/322025.12 | 74.7 | 66.7 | 95 | 89.4 | 71.4 | 93.1 | 79.8 | 31.1 | 72.1 | 64.3 | 80.6 | 78.2 | — | |
| CLIP-AdapterBackbone=ResNet-50, Shots=16, Blending ratio selection strategy=chosen per dataset (original paper results)2026.03 | 74.44 | 63.59 | 92.49 | 87.84 | 74.01 | 93.9 | 78.25 | 32.1 | 69.55 | 65.96 | — | 76.76 | 84.43 | |
| MaPLeModel=ViT-B/322025.12 | 74.1 | 66.7 | 95.1 | 91.7 | 66.9 | 89 | 82.1 | 28 | 72 | 63.4 | 83.3 | 77.3 | — | |
| CLIPClassifier=Prototype, Embedding Projection=Original, Backbone=ViT-B2026.03 | 73.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| CLIP-AdapterBackbone=ResNet-50, Shots=162026.03 | 73.35 | 59.02 | 92.28 | 84.92 | 73.49 | 94.56 | 73.96 | 34.19 | 68.14 | 65.7 | — | 77.3 | 83.24 | |
| PathCLIPBackbone=ResNet-50, Shots=16, Dual-view Vision Contrastive learning=reimplemented2026.03 | 73.35 | 63.5 | 92.7 | 84.9 | 73.8 | 94.1 | 77.5 | 34.3 | 70.3 | 66.4 | — | 79 | 81.9 | |
| KgCoOpModel=ViT-B/322025.12 | 72.1 | 65.4 | 94.4 | 90.8 | 67.3 | 86.1 | 81.7 | 23.7 | 71 | 65.1 | 70.1 | 77.5 | — | |
| CoCoOpModel=ViT-B/322025.12 | 70.7 | 66 | 94.3 | 91 | 64.6 | 82.5 | 81.9 | 22.6 | 69.8 | 59.7 | 70.4 | 75.3 | — | |
| TIP-AdapterBackbone=ResNet-50, Shots=162026.03 | 64.61 | 57.81 | 88.44 | 81.09 | 58.83 | 78.41 | 72.96 | 21.96 | 64 | 54.79 | — | 64.52 | 67.9 | |
| SVL-AdapterBackbone=ResNet-50, Shots=16, Semi-supervised learning (SSL)=false, Blending ratio selection strategy=zero-shot CLIP confidence scores2026.03 | 58.11 | 51.4 | 82.9 | 71.1 | 35 | 81.4 | 38.8 | 26.4 | 59.6 | 53.8 | — | 63.7 | 75.1 | |
| Zero-ShotBackbone=ResNet-502026.03 | 57.71 | 60.35 | 83.81 | 82.86 | 55.69 | 65.94 | 74.85 | 17.16 | 56.8 | 42.32 | — | 57.47 | 37.53 | |
| SVL-AdapterBackbone=ResNet-50, Shots=16, Blending ratio selection strategy=estimated2026.03 | — | — | 90 | 89.5 | 54 | 87 | 78 | 29.7 | 65.5 | 63 | — | 70.5 | 97.5 |