Image Classification on DTD (Accuracy)
94.3AccuracyMDA AM
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| MDA AMBackbone=ViT-B/32, Rotation Alignment=true2025.11 | 94.3 | — | |
| MDA AMBackbone=ViT-B/32, Rotation Alignment=false2025.11 | 92.9 | — | |
| MDA TABackbone=ViT-B/322025.11 | 91.5 | — | |
| TSV TABackbone=ViT-B/322025.11 | 90 | — | |
| ISO TABackbone=ViT-B/322025.11 | 87.1 | — | |
| theta_B fine-tuneModel=ViT-L/14, Mode=fine-tune2026.02 | 86.88 | — | |
| StableRepPre-training Data Source=Synthetic, Evaluation Protocol=5-way, 5-shot2023.06 | 86.4 | — | |
| StableRepPre-training dataset=CC12M, Pre-training data source=Syn, Evaluation protocol=Few-shot2023.06 | 86.4 | — | |
| StableRepPre-training dataset=RedCaps, Pre-training data source=Syn, Evaluation protocol=Few-shot2023.06 | 86.2 | — | |
| fine-tunedBackbone=ViT-L/142026.02 | 85.5 | — | |
| ISO-CLS TABackbone=ViT-B/322025.11 | 85.3 | — | |
| Fine-TunedBackbone=DINOv3 ViT-B/162025.12 | 84.5 | — | |
| PaRaMSModel Status=Protected Standalone (θ̂_def)2025.11 | 84.15 | — | |
| SynCLR + GMAILEvaluation Protocol=Zero-shot2026.02 | 83.67 | — | |
| ISO-CTS/LARVBackbone=ViT-L/142026.02 | 83.6 | 2.5 | |
| theta_B fine-tuneBackbone=ViT-L/142025.10 | 83.56 | — | |
| Linear ProbeBackbone=DINOv3 ViT-B/162025.12 | 83.5 | — | |
| CLIPPre-training Data Source=Real, Evaluation Protocol=5-way, 5-shot2023.06 | 83.3 | — | |
| CLIPPre-training dataset=CC12M, Pre-training data source=Real, Evaluation protocol=Few-shot2023.06 | 83.3 | — | |
| θBmode=fine-tune2026.02 | 83.29 | — | |
| θBMode=fine-tune2026.02 | 83.29 | — | |
| theta_A fine-tuneModel=ViT-B/16, Mode=fine-tune2026.02 | 83.29 | — | |
| theta_B fine-tuneBackbone=ViT-B/162025.10 | 83.19 | — | |
| ISO-C/LARVBackbone=ViT-L/142026.02 | 82.9 | 2.6 | |
| Task MatrixBackbone=DINOv3 ViT-B/16, best layer=112025.12 | 82.9 | — | |
| SimCLRPre-training dataset=RedCaps, Pre-training data source=Real, Evaluation protocol=Few-shot2023.06 | 82.7 | — | |
| CLIPPre-training dataset=RedCaps, Pre-training data source=Real, Evaluation protocol=Few-shot2023.06 | 82.6 | — | |
| CLIPPre-training dataset=RedCaps, Pre-training data source=Syn, Evaluation protocol=Few-shot2023.06 | 82.6 | — | |
| fine-tunedBackbone=ViT-B/162026.02 | 82.3 | — | |
| theta_A fine-tuneModel=laion2b, Protocol=fine-tune2026.02 | 82.24 | — | |
| SimCLRPre-training Data Source=Real, Evaluation Protocol=5-way, 5-shot2023.06 | 82.2 | — | |
| SimCLRPre-training dataset=CC12M, Pre-training data source=Real, Evaluation protocol=Few-shot2023.06 | 82.2 | — | |
| MergeGuardModel Status=Protected Standalone (θ̂_def)2025.11 | 82.16 | — | |
| CLIPPre-training Data Source=Synthetic, Evaluation Protocol=5-way, 5-shot2023.06 | 81.7 | — | |
| CLIPPre-training dataset=CC12M, Pre-training data source=Syn, Evaluation protocol=Few-shot2023.06 | 81.7 | — | |
| SynCLREvaluation Protocol=Zero-shot2026.02 | 79.9 | — | |
| SimCLRPre-training Data Source=Synthetic, Evaluation Protocol=5-way, 5-shot2023.06 | 79.8 | — | |
| SimCLRPre-training dataset=CC12M, Pre-training data source=Syn, Evaluation protocol=Few-shot2023.06 | 79.8 | — | |
| fine-tunedBackbone=ViT-B/322026.02 | 79.7 | — | |
| TSV-M/LARVBackbone=ViT-L/142026.02 | 79.5 | 4.2 | |
| SimCLRPre-training dataset=RedCaps, Pre-training data source=Syn, Evaluation protocol=Few-shot2023.06 | 79.5 | — | |
| Fine-Tuned*Note=Upper-bound2025.12 | 79.4 | — | |
| ADAProtection Status=None (theta_merge)2025.11 | 79.2 | — | |
| CLIP-ViT-B/16Img-text pairs=400M, Evaluation protocol=Linear probing2022.09 | 79.2 | — | |
| theta_B fine-tuneModel=laion400m, Protocol=fine-tune2026.02 | 77.81 | — | |
| MMFTModel=ViT, Method=Ours2026.01 | 77.68 | — | |
| Fine-TunedBackbone=CLIP ViT-B/32, Classes=472025.12 | 77.4 | — | |
| FLAVAImg-text pairs=70M, Evaluation protocol=Linear probing2022.09 | 77.3 | — | |
| Linear ProbeBackbone=CLIP ViT-B/32, Classes=472025.12 | 77.2 | — | |
| ISO-C/LARVBackbone=ViT-B/162026.02 | 76.8 | 4.8 | |
| FLYPModel=ViT, Method=FLYP2026.01 | 76.74 | — | |
| OmniVLImg-text pairs=14M*, Evaluation protocol=Linear probing2022.09 | 76.5 | — | |
| IsoCLIPIntra-modal=✓, Classifier=NCM, Backbone=ViT-B/16-open2026.03 | 75.9 | — | |
| Task MatrixBackbone=CLIP ViT-B/32, Best Layer=11, Classes=472025.12 | 75.7 | — | |
| TSV-M/LARVBackbone=ViT-B/162026.02 | 75.4 | 5.7 | |
| MMRLShots=16 shots2025.03 | 75.3 | — | |
| MMRLshots=162025.05 | 75.3 | — | |
| SimMatch-V2Backbone=ResNet-50 (4x), Pre-trained=ImageNet, Evaluation Protocol=Fine-tune2025.11 | 75.26 | — | |
| FTModel=ViT, Method=FT2026.01 | 74.72 | — | |
| ISO-CTS/LARVBackbone=ViT-B/162026.02 | 74.7 | 1 | |
| SupConBackbone=ResNet-50 (4x), Pre-trained=ImageNet, Evaluation Protocol=Fine-tune2025.11 | 74.6 | — | |
| BLIPImg-text pairs=14M, Evaluation protocol=Linear probing2022.09 | 74.6 | — | |
| ISO-CTS/LARVBackbone=ViT-B/322026.02 | 74.5 | 0.6 | |
| Image-ImageIntra-modal=✓, Classifier=NCM, Backbone=ViT-B/16-open2026.03 | 74.5 | — | |
| ISO-C/LARVBackbone=ViT-B/322026.02 | 74.4 | 9 | |
| MMRL++shots=162025.05 | 74.37 | — | |
| FLAVAImg-text pairs=14M, Evaluation protocol=Linear probing2022.09 | 74.2 | — | |
| CLOPBackbone=ResNet-50 (4x), Pre-trained=ImageNet, Evaluation Protocol=Fine-tune2025.11 | 73.85 | — | |
| MMAShots=16 shots2025.03 | 73.47 | — | |
| MMAshots=162025.05 | 73.47 | — | |
| IOTAShot=16-shot, Backbone=ViT-B/162026.01 | 73.4 | — | |
| ALBEFImg-text pairs=14M, Evaluation protocol=Linear probing2022.09 | 73.4 | — | |
| SimCLRBackbone=ResNet-50 (4x), Pre-trained=ImageNet, Evaluation Protocol=Fine-tune2025.11 | 73.2 | — | |
| SimMatchBackbone=ResNet-50 (4x), Pre-trained=ImageNet, Evaluation Protocol=Fine-tune2025.11 | 73.06 | — | |
| SsCLBackbone=ResNet-50 (4x), Pre-trained=ImageNet, Evaluation Protocol=Fine-tune2025.11 | 73 | — | |
| CBBudget (%)=100, Backbone=CLIP ViT-L/14-3362024.11 | 72.97 | — | |
| PromptSRCShots=16 shots2025.03 | 72.73 | — | |
| PromptSRCshots=162025.05 | 72.73 | — | |
| THESEUSModel=ViT-L/14, B (alignment batches)=102026.02 | 72.71 | — | |
| CCSSLBackbone=ResNet-50 (4x), Pre-trained=ImageNet, Evaluation Protocol=Fine-tune2025.11 | 72.7 | — | |
| FixMatchBackbone=ResNet-50 (4x), Pre-trained=ImageNet, Evaluation Protocol=Fine-tune2025.11 | 72.69 | — | |
| theta_B + delta_starBackbone=ViT-L/142025.10 | 72.66 | — | |
| PCBBudget (%)=100, Backbone=CLIP ViT-L/14-3362024.11 | 72.6 | — | |
| CS-AlignerBackbone=ViT-B/16, Number of Shots=16, Evaluation Protocol=fine-tuned2025.02 | 72.3 | — | |
| CLIP-AdapterBackbone=ViT-B/16, Shots=16, Blending ratio (alpha)=best2026.03 | 71.7 | — | |
| MMRLShots=8 shots2025.03 | 71.6 | — | |
| MMRLshots=82025.05 | 71.6 | — | |
| theta_B + delta_starBackbone=ViT-B/162025.10 | 71.44 | — | |
| MaPLeShots=16 shots2025.03 | 71.33 | — | |
| MaPLeshots=162025.05 | 71.33 | — | |
| ProTextShot=16-shot, Backbone=ViT-B/162026.01 | 71.01 | — | |
| TSV-M/LARVBackbone=ViT-B/322026.02 | 70.9 | 7 | |
| CLIP-AdapterBackbone=ViT-B/16, Number of Shots=16, Evaluation Protocol=fine-tuned2025.02 | 70.9 | — | |
| CBBudget (%)=75, Backbone=CLIP ViT-L/14-3362024.11 | 70.9 | — | |
| MMRL++shots=82025.05 | 70.83 | — | |
| HOSO-AdapterBackbone=ViT-B/16, Shots=16, Number of runs=32026.03 | 70.67 | — | |
| Linear probe CLIPShots=16 shots2025.03 | 69.96 | — | |
| Linear probe CLIPshots=162025.05 | 69.96 | — | |
| CoOpShots=16 shots2025.03 | 69.87 | — | |
| PromptSRCShots=8 shots2025.03 | 69.87 | — |