Image Classification on Stanford Cars (Adversarial Robustness)
68.91Top-1 Accuracy (Clean)CLBP
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| CLBPShots=16, Training Epochs=10, Training Attack (PGD)=2-step (1/255), Evaluation Attack (PGD-100)=100-step (1/255)2026.05 | 68.91 | 56.24 | — | |
| AdvMaPLeShots=16, Training Epochs=10, Training Attack (PGD)=2-step (1/255), Evaluation Attack (PGD-100)=100-step (1/255)2026.05 | 56.17 | 17.57 | — | |
| AdvVLPShots=16, Training Epochs=10, Training Attack (PGD)=2-step (1/255), Evaluation Attack (PGD-100)=100-step (1/255)2026.05 | 56 | 17.47 | — | |
| FAPShots=16, Training Epochs=10, Training Attack (PGD)=2-step (1/255), Evaluation Attack (PGD-100)=100-step (1/255)2026.05 | 54.23 | 19.23 | — | |
| CLIPBackbone=ViT-B/32, Evaluation Protocol=Zero-shot2026.03 | 54.1 | — | — | |
| WISE-FTBackbone=ViT-B/32, Pre-training Dataset=ImageNet, Evaluation Protocol=Zero-shot2026.03 | 53.18 | 0 | — | |
| GRACEBackbone=ViT-B/32, Pre-training Dataset=ImageNet, Evaluation Protocol=Zero-shot2026.03 | 52.1 | 1.9 | — | |
| CLIPYear=20212026.04 | 52.07 | — | — | |
| CLIPBackbone=ViT-B/32, Evaluation Protocol=Zero-shot2026.03 | 51.47 | 0 | — | |
| OursModel Configuration=5 trees2026.04 | 48.23 | — | — | |
| OursModel Configuration=1 tree2026.04 | 48.11 | — | — | |
| TPGMBackbone=ViT-B/32, Pre-training Dataset=ImageNet, Evaluation Protocol=Zero-shot2026.03 | 47.69 | 0 | — | |
| AoSYear=20252026.04 | 46.42 | — | — | |
| SPDBackbone=ViT-B/32, Pre-training Dataset=ImageNet, Evaluation Protocol=Zero-shot2026.03 | 46.31 | 0 | — | |
| AGFTBackbone=ViT-B/32, Evaluation Protocol=Zero-shot, Training Strategy=Adversarially fine-tuned2026.03 | 45.69 | — | — | |
| Vanilla FTBackbone=ViT-B/32, Pre-training Dataset=ImageNet, Evaluation Protocol=Zero-shot2026.03 | 45.44 | 0 | — | |
| FAREYear=20242026.04 | 45.06 | — | — | |
| GLADIATORBackbone=ViT-B/32, Evaluation Protocol=Zero-shot, Training Strategy=Adversarially fine-tuned2026.03 | 43.75 | — | — | |
| PMG-AFTBackbone=ViT-B/32, Evaluation Protocol=Zero-shot, Training Strategy=Adversarially fine-tuned2026.03 | 42.21 | — | — | |
| PMG-FTYear=20242026.04 | 40.49 | — | — | |
| TeCoABackbone=ViT-B/32, Evaluation Protocol=Zero-shot, Training Strategy=Adversarially fine-tuned2026.03 | 33.88 | — | — | |
| FAREBackbone=ViT-B/32, Pre-training Dataset=ImageNet, Evaluation Protocol=Zero-shot2026.03 | 33.87 | 1.3 | — | |
| FLYPBackbone=ViT-B/32, Pre-training Dataset=ImageNet, Evaluation Protocol=Zero-shot2026.03 | 33.72 | 0 | — | |
| PMG-AFTBackbone=ViT-B/32, Pre-training Dataset=ImageNet, Evaluation Protocol=Zero-shot2026.03 | 29.5 | 2 | — | |
| TGA-ZSRBackbone=ViT-B/32, Evaluation Protocol=Zero-shot, Training Strategy=Adversarially fine-tuned2026.03 | 29.34 | — | — | |
| LAATBackbone=ViT-B/32, Pre-training Dataset=ImageNet, Evaluation Protocol=Zero-shot2026.03 | 28.8 | 1.6 | — | |
| TeCoAYear=20232026.04 | 26.59 | — | — | |
| AdvVPShots=16, Training Epochs=10, Training Attack (PGD)=2-step (1/255), Evaluation Attack (PGD-100)=100-step (1/255)2026.05 | 14.83 | 3.57 | — | |
| TeCoABackbone=ViT-B/32, Pre-training Dataset=ImageNet, Evaluation Protocol=Zero-shot2026.03 | 10.68 | 1.5 | — | |
| A-TPTBackbone=ViT-B/162026.05 | — | 39.2 | — | |
| A-TPTBackbone=ViT-B/322026.05 | — | 31.8 | — | |
| CLIPBackbone=ViT-B/162026.05 | — | 0 | — | |
| CLIPBackbone=ViT-B/322026.05 | — | 0 | — | |
| GPT-4V2024.03 | — | — | 58.3 | |
| MTABackbone=ViT-B/162026.05 | — | 18.5 | — | |
| MTABackbone=ViT-B/322026.05 | — | 26.4 | — | |
| R-TPTBackbone=ViT-B/162026.05 | — | 34.7 | — | |
| R-TPTBackbone=ViT-B/322026.05 | — | 28.4 | — | |
| RARBase Model=LLaVA1.52024.03 | — | — | 72.6 | |
| RARBase Model=Intern-IXC22024.03 | — | — | 65.4 | |
| RARBase Model=Qwen-VL2024.03 | — | — | 81.6 | |
| TPT-EnsembleBackbone=ViT-B/162026.05 | — | 26 | — | |
| TPT-EnsembleBackbone=ViT-B/322026.05 | — | 25.9 | — | |
| TTCBackbone=ViT-B/162026.05 | — | 2.9 | — | |
| TTCBackbone=ViT-B/322026.05 | — | 2.3 | — |