Image Classification on 11 Natural Datasets (Transductive Setting)
77.8Average AccuracySOTA (CLIP-1+DINO(v2+v3))
Evaluation Results
| Method | Links | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| SOTA (CLIP-1+DINO(v2+v3))Visual Encoders=CLIP-1, DINOv2, and DINOv32025.06 | 77.8 | 79.9 | 72.8 | 34 | 87.2 | 78.5 | 90 | 95.3 | 84.5 | 96.6 | 55.9 | 81.1 | |
| SOTA (CLIP-1+DINOv3)Visual Encoders=CLIP-1 and DINOv32025.06 | 77.4 | 78.7 | 72.6 | 33 | 86.4 | 79.6 | 90.1 | 95.1 | 84.2 | 95.9 | 56.8 | 79.2 | |
| SOTA (CLIP-1+DINOv2)Visual Encoders=CLIP-1 and DINOv22025.06 | 75.7 | 77.7 | 71.9 | 28 | 84.4 | 73.9 | 89 | 94.5 | 83.9 | 96.1 | 52.8 | 80.3 | |
| GTA-CLIPSetting=Transductive2025.06 | 74.5 | 71.9 | 73.5 | 29.3 | 76.4 | 72.1 | 87.4 | 93.4 | 82.1 | 95.5 | 58.5 | 79.1 | |
| SOTA (CLIP-1+CLIP-2)Visual Encoders=CLIP-1 (ViT-B/16) and CLIP-2 (RN50)2025.06 | 72.5 | 69.9 | 70.2 | 22.6 | 82.9 | 69.9 | 86 | 92.4 | 75.1 | 92.9 | 56 | 79.4 | |
| ADAPTSetting=Transductive2025.06 | 72.4 | 71.6 | 72.3 | 30.8 | 65.9 | 71.3 | 85.1 | 92.6 | 80.1 | 95.5 | 56.9 | 73.9 | |
| ECALPSetting=Transductive2025.06 | 70.5 | 71.3 | 70.4 | 29.5 | 56.5 | 68.2 | 85.7 | 92.3 | 76 | 94.4 | 56.3 | 75.4 | |
| TransCLIP-2Setting=Transductive2025.06 | 70.3 | 70.4 | 68.9 | 26.9 | 66.1 | 69.5 | 87.1 | 92.5 | 76.5 | 92.7 | 48.6 | 74.1 | |
| Stat.ASetting=Transductive2025.06 | 69.9 | 69.9 | 68.7 | 24.7 | 67.3 | 68 | 87.1 | 92.4 | 75.2 | 94.2 | 48.4 | 73.5 | |
| GDASetting=Transductive2025.06 | 67.6 | 67.3 | 63.9 | 25.5 | 59.5 | 67 | 86.4 | 90.8 | 74 | 93.8 | 45.3 | 70.3 | |
| ZLaPSetting=Transductive2025.06 | 67.5 | 69.7 | 67.8 | 26.3 | 57.7 | 66.8 | 87.2 | 87.9 | 67.9 | 91.8 | 46 | 73.8 | |
| CLIP-1Visual Encoder=ViT-B/162025.06 | 65.2 | 66.6 | 62.5 | 24.7 | 48.3 | 65.6 | 85.9 | 89.1 | 70.7 | 93.2 | 43.5 | 67.5 | |
| CLIP-2Visual Encoder=RN502025.06 | 56.1 | 58.2 | 58.8 | 15.7 | 23.7 | 55.7 | 74 | 83.6 | 61.8 | 85.9 | 40.4 | 58.8 |