Image Classification on DTD (Describable Textures Dataset)
79.2AccuracySAE-FT
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SAE-FTBackbone=ViT-B/162026.05 | 79.2 | — | |
| L2 regularizationBackbone=ViT-B/162026.05 | 78.99 | — | |
| LeJEPA*Backbone=ViT-L/14, Pre-training Epochs=100, Training Heuristics=w/o heuristics, Linear Probe Training Epochs=1002026.06 | 78.3 | — | |
| Standard fine-tuningBackbone=ViT-B/162026.05 | 77.82 | — | |
| VISRegBackbone=ViT-L/14, Pre-training Epochs=400, Training Heuristics=w/o heuristics, Linear Probe Training Epochs=102026.06 | 76.5 | — | |
| VISRegBackbone=ViT-L/14, Pre-training Epochs=100, Training Heuristics=w/o heuristics, Linear Probe Training Epochs=1002026.06 | 76.3 | — | |
| VISRegBackbone=ViT-B/16, Pre-training Epochs=400, Training Heuristics=w/o heuristics, Linear Probe Training Epochs=102026.06 | 75.7 | — | |
| CLIP-LoRA + FMAShots=162025.10 | 75.4 | — | |
| iBOTBackbone=ViT-L/16, Pre-training Epochs=250, Training Heuristics=w/ heuristics, Linear Probe Training Epochs=102026.06 | 75.3 | — | |
| DINOBackbone=ViT-B/16, Pre-training Epochs=400, Training Heuristics=w/ heuristics, Linear Probe Training Epochs=102026.06 | 74.3 | — | |
| iBOTBackbone=ViT-B/16, Pre-training Epochs=400, Training Heuristics=w/ heuristics, Linear Probe Training Epochs=102026.06 | 74.1 | — | |
| MoCoV3Backbone=ViT-B/16, Pre-training Epochs=300, Training Heuristics=w/ heuristics, Linear Probe Training Epochs=102026.06 | 73.7 | — | |
| TOGAShots=16, Venue=-2026.03 | 73.6 | — | |
| CLIP-LoRAShots=162025.10 | 73 | — | |
| MAEBackbone=ViT-L/16, Pre-training Epochs=1600, Training Heuristics=w/o heuristics, Linear Probe Training Epochs=102026.06 | 72.8 | — | |
| PLOT++Shots=162025.10 | 71.4 | — | |
| TIP-AdapterShots=162025.10 | 70.8 | — | |
| ViT-BAdaptation Strategy=Fine-tuning, Params=85.8M, Pre-training Dataset=ImageNet-1K2026.05 | 70.8 | — | |
| CoOpShots=162025.10 | 70 | — | |
| ViT-BAdaptation Strategy=LoRA, FFN, Params=738K, Pre-training Dataset=ImageNet-1K2026.05 | 70 | — | |
| I-JEPABackbone=ViT-H/14, Pre-training Epochs=300, Training Heuristics=w/ heuristics, Linear Probe Training Epochs=102026.06 | 69.9 | — | |
| data2vecBackbone=ViT-L/14, Pre-training Epochs=1600, Training Heuristics=w/ heuristics, Linear Probe Training Epochs=102026.06 | 69.7 | — | |
| TOGAShots=8, Venue=-2026.03 | 69.6 | — | |
| ViT-BAdaptation Strategy=LoRA, QV, Params=296K, Pre-training Dataset=ImageNet-1K2026.05 | 69.5 | — | |
| bViT-BAdaptation Strategy=LoRA, FFN, Params=62K, Pre-training Dataset=ImageNet-1K2026.05 | 69.5 | — | |
| bViT-BAdaptation Strategy=Fine-tuning, Params=7.8M, Pre-training Dataset=ImageNet-1K2026.05 | 69.3 | — | |
| bViT-BAdaptation Strategy=LoRA, QV, Params=25K, Pre-training Dataset=ImageNet-1K2026.05 | 69 | — | |
| ProGradShots=162025.10 | 68.8 | — | |
| KgCoOpShots=162025.10 | 68.7 | — | |
| bViT-BAdaptation Strategy=Time embedding tuning, Params=9K, Pre-training Dataset=ImageNet-1K2026.05 | 67.9 | — | |
| CLIP-LoRA + FMAShots=42025.10 | 67 | — | |
| Align+MixShot=16, Backbone=ResNet-502026.03 | 67 | — | |
| GDAShot=16, Backbone=ResNet-502026.03 | 66.1 | — | |
| ViT-BAdaptation Strategy=Linear probing, Params=—, Pre-training Dataset=ImageNet-1K2026.05 | 66.1 | — | |
| CoCoOpShots=162025.10 | 65.8 | — | |
| bViT-BAdaptation Strategy=Linear probing, Params=—, Pre-training Dataset=ImageNet-1K2026.05 | 65.6 | — | |
| TOGAShots=4, Venue=-2026.03 | 64.5 | — | |
| CLIP-LoRAShots=42025.10 | 64 | — | |
| TIP-XShot=16, Backbone=ResNet-502026.03 | 63.5 | — | |
| PLOT++Shots=42025.10 | 62.4 | — | |
| GRIPModel=Idefics2-8B2026.06 | 61.3 | — | |
| TIP-AdapterShot=16, Backbone=ResNet-502026.03 | 60.9 | — | |
| TIP-AdapterShots=42025.10 | 59.8 | — | |
| ProGradShots=42025.10 | 59.7 | — | |
| CoOpShots=42025.10 | 59.5 | — | |
| CLIP-AdapterShots=162025.10 | 59.4 | — | |
| Align+MixShot=4, Backbone=ResNet-502026.03 | 58.8 | — | |
| KgCoOpShots=42025.10 | 58.7 | — | |
| TOGAShots=2, Venue=-2026.03 | 58.4 | — | |
| GDAShot=4, Backbone=ResNet-502026.03 | 57.4 | — | |
| ITOBackbone=ViT-L/16, Pre-training Dataset=DataComp-1B, Training Epochs=1, Evaluation Protocol=Zero-shot2026.03 | 56 | — | |
| CLIPModel=Idefics2-8B2026.06 | 55.8 | — | |
| CoCoOpShots=42025.10 | 55.7 | — | |
| TOGAShots=1, Venue=-2026.03 | 55.2 | — | |
| TIP-XShot=4, Backbone=ResNet-502026.03 | 55.2 | — | |
| CLIP-LoRA + FMAShots=12025.10 | 55.1 | — | |
| ViTModel=Idefics2-8B2026.06 | 54.8 | — | |
| PLOT++Shots=12025.10 | 54.6 | — | |
| CLIP-LoRAShots=12025.10 | 54.1 | — | |
| DINOModel=Idefics2-8B2026.06 | 54.1 | — | |
| TIP-AdapterShot=4, Backbone=ResNet-502026.03 | 54 | — | |
| ITO sub2Backbone=ViT-L/16, Pre-training Dataset=DataComp-1B, Training Epochs=1, Evaluation Protocol=Zero-shot2026.03 | 53.9 | — | |
| GRIPModel=Qwen2.5-VL-7B2026.06 | 53 | — | |
| ProGradShots=12025.10 | 52.8 | — | |
| KgCoOpShots=12025.10 | 52.7 | — | |
| CoCoOpShots=12025.10 | 52.6 | — | |
| DINOModel=Qwen2.5-VL-7B2026.06 | 52.3 | — | |
| ViTModel=Qwen2.5-VL-7B2026.06 | 52.2 | — | |
| TIP-AdapterShots=12025.10 | 51.6 | — | |
| CLIPModel=Qwen2.5-VL-7B2026.06 | 51.5 | — | |
| CLIPBackbone=ViT-L/16, Pre-training Dataset=DataComp-1B, Training Epochs=1, Evaluation Protocol=Zero-shot2026.03 | 51.1 | — | |
| CoOpShots=12025.10 | 50.1 | — | |
| Zero ShotModel=Qwen2.5-VL-7B2026.06 | 48.7 | — | |
| Align+MixShot=1, Backbone=ResNet-502026.03 | 48.3 | — | |
| TIP-XShot=1, Backbone=ResNet-502026.03 | 47 | — | |
| TIP-AdapterShot=1, Backbone=ResNet-502026.03 | 46.2 | — | |
| CLIP-AdapterShots=42025.10 | 46.1 | — | |
| GDAShot=1, Backbone=ResNet-502026.03 | 46.1 | — | |
| CLIP-AdapterShots=12025.10 | 44.2 | — | |
| CLIPShots=02025.10 | 43.8 | — | |
| Zero ShotModel=Idefics2-8B2026.06 | 43.7 | — | |
| CALIPShot=0, Backbone=ResNet-502026.03 | 42.4 | — | |
| CLIPShot=0, Backbone=ResNet-502026.03 | 42.3 | — | |
| CLIPShots=0, Venue=ICML’222026.03 | 42.2 | — | |
| SigLIPPre-training Dataset=CC12M, Backbone=ViT-B/16, Epochs=30, Evaluation protocol=Zero-shot2026.03 | 27.1 | — | |
| ITO sub2Pre-training Dataset=CC12M, Backbone=ViT-B/16, Epochs=30, Evaluation protocol=Zero-shot2026.03 | 25.5 | — | |
| ITOPre-training Dataset=CC12M, Backbone=ViT-B/16, Epochs=30, Evaluation protocol=Zero-shot2026.03 | 23.7 | — | |
| SLIPPre-training Dataset=CC12M, Backbone=ViT-B/16, Epochs=30, Evaluation protocol=Zero-shot2026.03 | 22.6 | — | |
| FLAIRPre-training Dataset=CC12M, Backbone=ViT-B/16, Epochs=30, Evaluation protocol=Zero-shot2026.03 | 19.4 | — | |
| CLIPPre-training Dataset=CC12M, Backbone=ViT-B/16, Epochs=30, Evaluation protocol=Zero-shot2026.03 | 17.6 | — | |
| BARShots=16, Model Access Type=Black-box2024.07 | — | 47 | |
| BlackVIPShots=16, Model Access Type=Black-box2024.07 | — | 45.2 | |
| BlackVIP-SEShots=16, Model Access Type=Black-box2024.07 | — | 45.1 | |
| CEEvaluation Protocol=Linear Probing2026.05 | — | 69.9 | |
| ETF + DREvaluation Protocol=Linear Probing2026.05 | — | 64 | |
| NONLEvaluation Protocol=Linear Probing2026.05 | — | 69.4 | |
| NormFaceEvaluation Protocol=Linear Probing2026.05 | — | 70.1 | |
| NTCEEvaluation Protocol=Linear Probing2026.05 | — | 70 | |
| VP (white-box)Shots=16, Model Access Type=White-box2024.07 | — | 61.9 | |
| VP w/ SPSA-GCShots=16, Model Access Type=Black-box2024.07 | — | 44.5 |