Fine-grained classification on DTD (Adversarial Robustness)
54Clean AccuracyR-TPT
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| R-TPTBackbone=CLIP-ViT-L/14, Perturbation budget (epsilon)=4.02025.12 | 54 | 38 | |
| R-TPTBackbone=ViT-L/14, Epsilon=4/2552026.05 | 54 | 38 | |
| MTABackbone=CLIP-ViT-L/14, Perturbation budget (epsilon)=4.02025.12 | 53.4 | 27.2 | |
| MTABackbone=ViT-L/14, Epsilon=4/2552026.05 | 53.4 | 27.2 | |
| CLIPBackbone=CLIP-ViT-L/14, Perturbation budget (epsilon)=4.02025.12 | 52.4 | 0 | |
| CLIPBackbone=ViT-L/14, Epsilon=4/2552026.05 | 52.4 | 0 | |
| TTPBackbone=CLIP-ViT-L/14, Perturbation budget (epsilon)=4.02025.12 | 52.3 | 41.3 | |
| TTPBackbone=ViT-L/14, Epsilon=4/2552026.05 | 52.3 | 41.3 | |
| AGCBackbone=ViT-L/14, Epsilon=4/2552026.05 | 52.2 | 95.2 | |
| EnsembleBackbone=CLIP-ViT-L/14, Perturbation budget (epsilon)=4.02025.12 | 51.3 | 31.3 | |
| EnsembleBackbone=ViT-L/14, Epsilon=4/2552026.05 | 50.2 | 38.7 | |
| TTCBackbone=CLIP-ViT-L/14, Perturbation budget (epsilon)=4.02025.12 | 49.7 | 6.2 | |
| TTCBackbone=ViT-L/14, Epsilon=4/2552026.05 | 49.7 | 6.2 | |
| MTABackbone=CLIP-ViT-B/16, Perturbation bound (epsilon)=4.02025.12 | 46.5 | 16.2 | |
| R-TPTBackbone=CLIP-ViT-B/16, Perturbation bound (epsilon)=4.02025.12 | 46.4 | 32.8 | |
| TPTBackbone=CLIP-ViT-B/16, Augmented views=15, Perturbation (epsilon)=4.02026.06 | 46.2 | 0 | |
| SS-TPTBackbone=CLIP-ViT-B/16, Augmented views=15, Perturbation (epsilon)=4.02026.06 | 45 | 31.9 | |
| R-TPTBackbone=CLIP-ViT-B/16, Augmented views=15, Perturbation (epsilon)=4.02026.06 | 44.9 | 26.7 | |
| CLIPBackbone=CLIP-ViT-B/16, Perturbation bound (epsilon)=4.02025.12 | 44.4 | 0 | |
| CLIPBackbone=CLIP-ViT-B/16, Augmented views=15, Perturbation (epsilon)=4.02026.06 | 44.4 | 0 | |
| TTPBackbone=CLIP-ViT-B/16, Perturbation bound (epsilon)=4.02025.12 | 44.1 | 36 | |
| EnsembleBackbone=CLIP-ViT-B/16, Augmented views=15, Perturbation (epsilon)=4.02026.06 | 44 | 17.4 | |
| MTABackbone=CLIP-ViT-B/32, Perturbation Bound (epsilon)=4.02025.12 | 43.8 | 28.8 | |
| TAPTBackbone=CLIP-ViT-B/16, Augmented views=15, Perturbation (epsilon)=4.02026.06 | 43.6 | 12.6 | |
| EnsembleBackbone=CLIP-ViT-B/16, Perturbation bound (epsilon)=4.02025.12 | 43.2 | 25.1 | |
| CLIPBackbone=CLIP-ViT-B/32, Perturbation Bound (epsilon)=4.02025.12 | 43 | 0 | |
| TTPBackbone=CLIP-ViT-B/32, Perturbation Bound (epsilon)=4.02025.12 | 42.8 | 32.2 | |
| TPTBackbone=CLIP-ResNet50, Epsilon=1.02025.04 | 42.4 | 4.3 | |
| R-TPTBackbone=CLIP-ViT-B/32, Perturbation Bound (epsilon)=4.02025.12 | 42.1 | 29.1 | |
| DOCBackbone=CLIP-ViT-B/16, Augmented views=15, Perturbation (epsilon)=4.02026.06 | 42 | 12.2 | |
| TTCBackbone=CLIP-ViT-B/16, Augmented views=15, Perturbation (epsilon)=4.02026.06 | 41.8 | 5.5 | |
| C-TPTBackbone=CLIP-ResNet50, Epsilon=1.02025.04 | 41.5 | 1.3 | |
| R-TPTBackbone=CLIP-ResNet50, Epsilon=1.02025.04 | 41.3 | 33.5 | |
| TTCBackbone=CLIP-ViT-B/16, Perturbation bound (epsilon)=4.02025.12 | 41 | 4.5 | |
| CLIPBackbone=CLIP-ResNet50, Epsilon=1.02025.04 | 40.4 | 0.8 | |
| MTABackbone=CLIP-ResNet50, Epsilon=1.02025.04 | 40.3 | 18.8 | |
| EnsembleBackbone=CLIP-ViT-B/32, Perturbation Bound (epsilon)=4.02025.12 | 39.8 | 28.6 | |
| TTCBackbone=CLIP-ViT-B/32, Perturbation Bound (epsilon)=4.02025.12 | 37.3 | 4.7 | |
| EnsembleBackbone=CLIP-ResNet50, Epsilon=1.02025.04 | 37.1 | 29.5 | |
| R-TPT + ET3Pre-trained Model=TeCoA CLIP-ViT-B/32, Attack Protocol=PGD-100 (ϵ = 4/255), Defense Category=Multiple Augmentations + Prompt Tuning, Input Modality=Vision + Text2026.03 | 26.42 | 19.92 | |
| C-TPT*Pre-trained Model=TeCoA CLIP-ViT-B/32, Attack Protocol=PGD-100 (ϵ = 4/255), Defense Category=Multiple Augmentations + Prompt Tuning, Input Modality=Vision + Text2026.03 | 26.2 | 12.4 | |
| R-TPTPre-trained Model=TeCoA CLIP-ViT-B/32, Attack Protocol=PGD-100 (ϵ = 4/255), Defense Category=Multiple Augmentations + Prompt Tuning, Input Modality=Vision + Text2026.03 | 25.59 | 18.09 | |
| TPT*Pre-trained Model=TeCoA CLIP-ViT-B/32, Attack Protocol=PGD-100 (ϵ = 4/255), Defense Category=Multiple Augmentations + Prompt Tuning, Input Modality=Vision + Text2026.03 | 25.5 | 14.6 | |
| SS-TPTBackbone=CLIP-ViT-B/32, Pre-trained=TeCoA, Augmented views=15, Epsilon=4.02026.06 | 25 | 17.7 | |
| TPTBackbone=CLIP-ViT-B/32, Pre-trained=TeCoA, Augmented views=15, Epsilon=4.02026.06 | 24.9 | 14.3 | |
| CLIP (Robust)Pre-trained Model=TeCoA CLIP-ViT-B/32, Attack Protocol=PGD-100 (ϵ = 4/255), Defense Category=Baseline2026.03 | 24.53 | 10.7 | |
| CLIPBackbone=CLIP-ViT-B/32, Pre-trained=TeCoA, Augmented views=15, Epsilon=4.02026.06 | 24.5 | 11.8 | |
| MTA*Pre-trained Model=TeCoA CLIP-ViT-B/32, Attack Protocol=PGD-100 (ϵ = 4/255), Defense Category=Multiple Augmentations, Input Modality=Vision Only2026.03 | 24.4 | 13.5 | |
| R-TPTBackbone=CLIP-ViT-B/32, Pre-trained=TeCoA, Augmented views=15, Epsilon=4.02026.06 | 24.4 | 17 | |
| ET3Pre-trained Model=TeCoA CLIP-ViT-B/32, Attack Protocol=PGD-100 (ϵ = 4/255), Defense Category=Lightweight Defense, Input Modality=Vision Only2026.03 | 24 | 13.18 | |
| TTCPre-trained Model=TeCoA CLIP-ViT-B/32, Attack Protocol=PGD-100 (ϵ = 4/255), Defense Category=Lightweight Defense, Input Modality=Vision Only2026.03 | 23.99 | 11.38 | |
| EnsemblePre-trained Model=TeCoA CLIP-ViT-B/32, Attack Protocol=PGD-100 (ϵ = 4/255), Defense Category=Multiple Augmentations, Input Modality=Vision Only2026.03 | 23.82 | 15.96 | |
| EnsembleBackbone=CLIP-ViT-B/32, Pre-trained=TeCoA, Augmented views=15, Epsilon=4.02026.06 | 23.3 | 16 | |
| Ensemble + ET3Pre-trained Model=TeCoA CLIP-ViT-B/32, Attack Protocol=PGD-100 (ϵ = 4/255), Defense Category=Multiple Augmentations, Input Modality=Vision Only2026.03 | 22.64 | 17.55 |