Fine-grained Classification on Pets
93.7AccuracyMTA
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| MTABackbone=ViT-L/14, Epsilon=4/2552026.05 | 93.7 | 64.9 | |
| R-TPTBackbone=ViT-L/14, Epsilon=4/2552026.05 | 93.7 | 72.9 | |
| R-TPTCategory=Tuning-based, Backbone=CLIP-ViT-L/14, Attack Protocol=PGD-100, eps=4.02026.06 | 93.7 | 72.9 | |
| MTACategory=Tuning-free, Backbone=CLIP-ViT-L/14, Attack Protocol=PGD-100, eps=4.02026.06 | 93.7 | 64.9 | |
| MACCategory=Tuning-free, Backbone=CLIP-ViT-L/14, Attack Protocol=PGD-100, eps=4.02026.06 | 93.6 | 87 | |
| AGCBackbone=ViT-L/14, Epsilon=4/2552026.05 | 93.4 | 98.5 | |
| CLIPBackbone=ViT-L/14, Epsilon=4/2552026.05 | 93.1 | 0 | |
| TTPBackbone=ViT-L/14, Epsilon=4/2552026.05 | 93.1 | 76.3 | |
| CLIPCategory=Baseline, Backbone=CLIP-ViT-L/14, Attack Protocol=PGD-100, eps=4.02026.06 | 93.1 | 0 | |
| EnsembleBackbone=ViT-L/14, Epsilon=4/2552026.05 | 92.5 | 77.6 | |
| TTCBackbone=ViT-L/14, Epsilon=4/2552026.05 | 92.2 | 7.6 | |
| TTCCategory=Tuning-free, Backbone=CLIP-ViT-L/14, Attack Protocol=PGD-100, eps=4.02026.06 | 92.2 | 6.4 | |
| CLIPBackbone=ViT-B/16, Pre-train=CLIP, Epsilon (ϵ)=4/255, Attack Protocol=PGD-1002026.05 | 88.3 | 0 | |
| TTPBackbone=ViT-B/16, Pre-train=CLIP, Epsilon (ϵ)=4/255, Attack Protocol=PGD-1002026.05 | 88.3 | 64.7 | |
| CLIPmode=Zero-shot2026.02 | 88.25 | — | |
| C-TPTmode=Test-Time Adaptation2026.02 | 88.14 | — | |
| MTABackbone=ViT-B/16, Pre-train=CLIP, Epsilon (ϵ)=4/255, Attack Protocol=PGD-1002026.05 | 88 | 51.8 | |
| MTAmode=Test-Time Adaptation2026.02 | 87.9 | — | |
| TPSmode=Test-Time Adaptation2026.02 | 87.44 | — | |
| ZEROmode=Test-Time Adaptation2026.02 | 87.33 | — | |
| TPTmode=Test-Time Adaptation2026.02 | 87.22 | — | |
| TTLmode=Test-Time Adaptation2026.02 | 87.22 | — | |
| R-TPTBackbone=ViT-B/16, Pre-train=CLIP, Epsilon (ϵ)=4/255, Attack Protocol=PGD-1002026.05 | 87.2 | 60.2 | |
| AGCBackbone=ViT-B/16, Pre-train=CLIP, Epsilon (ϵ)=4/255, Attack Protocol=PGD-1002026.05 | 87 | 96.1 | |
| RLCFmode=Test-Time Adaptation2026.02 | 86.97 | — | |
| R-TPTmode=Test-Time Adaptation2026.02 | 86.73 | — | |
| Fair Context Learning (FCL)mode=Test-Time Adaptation2026.02 | 86.54 | — | |
| CLIPBackbone=CLIP2022.11 | 85.4 | — | |
| EnsembleBackbone=ViT-B/16, Pre-train=CLIP, Epsilon (ϵ)=4/255, Attack Protocol=PGD-1002026.05 | 85 | 65.9 | |
| R-TPTBackbone=ResNet50, Attack Type=PGD, Epsilon (ε)=1.02026.05 | 84.6 | 74.2 | |
| AGCBackbone=ResNet50, Attack Type=PGD, Epsilon (ε)=1.02026.05 | 84 | 85.1 | |
| CLIPBackbone=ResNet50, Attack Type=PGD, Epsilon (ε)=1.02026.05 | 83.6 | 0 | |
| TTPBackbone=ResNet50, Attack Type=PGD, Epsilon (ε)=1.02026.05 | 83.5 | 70.3 | |
| TTCBackbone=ViT-B/16, Pre-train=CLIP, Epsilon (ϵ)=4/255, Attack Protocol=PGD-1002026.05 | 82.3 | 10.4 | |
| APT+TeCoABackbone=ResNet50, Attack Type=PGD, Epsilon (ε)=1.02026.05 | 79.3 | 79 | |
| TeCoABackbone=ResNet50, Attack Type=PGD, Epsilon (ε)=1.02026.05 | 76 | 75.8 | |
| GIF-SDBackbone=ResNet-50, Expansion Ratio=30x2022.11 | 73.4 | — | |
| CLIP (Robust)Pre-trained Model=TeCoA CLIP-ViT-B/32, Attack Protocol=PGD-100 (ϵ = 4/255), Defense Category=Baseline2026.03 | 66.88 | 15.94 | |
| ET3Pre-trained Model=TeCoA CLIP-ViT-B/32, Attack Protocol=PGD-100 (ϵ = 4/255), Defense Category=Lightweight Defense, Input Modality=Vision Only2026.03 | 66.86 | 27.15 | |
| GIF-DALLEBackbone=ResNet-50, Expansion Ratio=30x2022.11 | 66.4 | — | |
| MTA*Pre-trained Model=TeCoA CLIP-ViT-B/32, Attack Protocol=PGD-100 (ϵ = 4/255), Defense Category=Multiple Augmentations, Input Modality=Vision Only2026.03 | 66.2 | 31.2 | |
| C-TPT*Pre-trained Model=TeCoA CLIP-ViT-B/32, Attack Protocol=PGD-100 (ϵ = 4/255), Defense Category=Multiple Augmentations + Prompt Tuning, Input Modality=Vision + Text2026.03 | 66.1 | 19.5 | |
| TPT*Pre-trained Model=TeCoA CLIP-ViT-B/32, Attack Protocol=PGD-100 (ϵ = 4/255), Defense Category=Multiple Augmentations + Prompt Tuning, Input Modality=Vision + Text2026.03 | 65.2 | 27.4 | |
| TTCPre-trained Model=TeCoA CLIP-ViT-B/32, Attack Protocol=PGD-100 (ϵ = 4/255), Defense Category=Lightweight Defense, Input Modality=Vision Only2026.03 | 65 | 21.1 | |
| R-TPT + ET3Pre-trained Model=TeCoA CLIP-ViT-B/32, Attack Protocol=PGD-100 (ϵ = 4/255), Defense Category=Multiple Augmentations + Prompt Tuning, Input Modality=Vision + Text2026.03 | 63.86 | 45.71 | |
| R-TPTPre-trained Model=TeCoA CLIP-ViT-B/32, Attack Protocol=PGD-100 (ϵ = 4/255), Defense Category=Multiple Augmentations + Prompt Tuning, Input Modality=Vision + Text2026.03 | 63.75 | 41.43 | |
| Ensemble + ET3Pre-trained Model=TeCoA CLIP-ViT-B/32, Attack Protocol=PGD-100 (ϵ = 4/255), Defense Category=Multiple Augmentations, Input Modality=Vision Only2026.03 | 62.03 | 43.17 | |
| DALL-E2Backbone=ResNet-50, Expansion Ratio=30x2022.11 | 61.7 | — | |
| EnsemblePre-trained Model=TeCoA CLIP-ViT-B/32, Attack Protocol=PGD-100 (ϵ = 4/255), Defense Category=Multiple Augmentations, Input Modality=Vision Only2026.03 | 59.96 | 38.35 | |
| SDBackbone=ResNet-50, Expansion Ratio=30x2022.11 | 57.9 | — | |
| GIF-MAEBackbone=ResNet-50, Expansion Ratio=30x2022.11 | 52.4 | — | |
| RandAugmentBackbone=ResNet-50, Expansion Ratio=30x2022.11 | 48 | — | |
| MAEBackbone=ResNet-50, Expansion Ratio=30x2022.11 | 39.9 | — | |
| CutoutBackbone=ResNet-50, Expansion Ratio=30x2022.11 | 38.7 | — | |
| GridMaskBackbone=ResNet-50, Expansion Ratio=30x2022.11 | 37.6 | — | |
| APTBackbone=ResNet50, Attack Type=PGD, Epsilon (ε)=1.02026.05 | 31.9 | 3.8 | |
| Distillation of CLIPBackbone=ResNet-502022.11 | 11.1 | — | |
| OriginalBackbone=ResNet-50, Expansion Ratio=1x2022.11 | 6.8 | — |