Image Classification on FGVC Aircraft
88.01AccuracySFDA
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| SFDATop-k=k=32022.07 | 88.01 | — | — | |
| SFDAcomTop-k=k=32022.07 | 88.01 | — | — | |
| SFDATop-k=k=22022.07 | 87.71 | — | — | |
| SFDAcomTop-k=k=22022.07 | 87.71 | — | — | |
| LogMETop-k=k=22022.07 | 87.23 | — | — | |
| LogMETop-k=k=32022.07 | 87.23 | — | — | |
| NLEEPTop-k=k=32022.07 | 86.98 | — | — | |
| LaCLIPArchitecture=ViT-B/16, Pre-training Data=LAION-400M, Protocol=5-way 5-shot2023.05 | 86.1 | — | — | |
| NLEEPTop-k=k=22022.07 | 85.17 | — | — | |
| LaCLIPArchitecture=ViT-B/32, Pre-training Data=LAION-400M, Protocol=5-way 5-shot2023.05 | 82.2 | — | — | |
| CLIPArchitecture=ViT-B/16, Pre-training Data=LAION-400M, Protocol=5-way 5-shot2023.05 | 80.8 | — | — | |
| SVD-ViTComponents=SPC, Layer=112026.02 | 80.68 | — | — | |
| SVD-ViTComponents=SPC + ID-RSVD, Layer=112026.02 | 79.93 | — | — | |
| SVD-ViTComponents=SPC + SSVA, Layer=102026.02 | 79.84 | — | — | |
| SVD-ViTComponents=SPC + SSVA, Layer=112026.02 | 79.63 | — | — | |
| MLCDPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Protocol=Linear Probe2024.07 | 79.6 | — | — | |
| SVD-ViTComponents=SPC + SSVA, Layer=92026.02 | 79.45 | — | — | |
| SVD-ViTComponents=SPC + ID-RSVD, Layer=92026.02 | 79.45 | — | — | |
| SVD-ViTComponents=SPC, Layer=102026.02 | 79.27 | — | — | |
| CLIPArchitecture=ViT-B/32, Pre-training Data=LAION-400M, Protocol=5-way 5-shot2023.05 | 78.9 | — | — | |
| SVD-ViTComponents=SPC + SSVA + ID-RSVD, Layer=102026.02 | 78.82 | — | — | |
| SVD-ViTComponents=SPC, Layer=92026.02 | 78.61 | — | — | |
| SVD-ViTComponents=SPC + ID-RSVD, Layer=102026.02 | 78.46 | — | — | |
| SVD-ViTComponents=SPC + SSVA + ID-RSVD, Layer=112026.02 | 78.1 | — | — | |
| ViTCLS=1, Layer=-2026.02 | 77.86 | — | — | |
| SVD-ViTComponents=SPC + SSVA + ID-RSVD, Layer=92026.02 | 77.74 | — | — | |
| ViTCLS=8, Layer=-2026.02 | 76.27 | — | — | |
| UNICOMPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Protocol=Linear Probe2024.07 | 74.5 | — | — | |
| CLIPPre-training Data=WIT-400M, Backbone=ViT-L/14, Evaluation Protocol=Linear Probe2024.07 | 69.4 | — | — | |
| LaCLIPArchitecture=ViT-B/16, Pre-training Data=CC12M, Protocol=5-way 5-shot2023.05 | 68.1 | — | — | |
| CLIP+Pre-training Data=WIT-400M, Backbone=ViT-L/14, Evaluation Protocol=Linear Probe2024.07 | 68 | — | — | |
| OpenCLIPPre-training Data=LAION-400M, Backbone=ViT-L/14, Evaluation Protocol=Linear Probe2024.07 | 64.4 | — | — | |
| CLIPArchitecture=ViT-B/16, Pre-training Data=CC12M, Protocol=5-way 5-shot2023.05 | 62 | — | — | |
| EsViTBackbone=Swin-T, Evaluation Protocol=Linear probe2021.06 | 61.1 | — | — | |
| RényiCLEvaluation Protocol=linear evaluation, Multi-crop Augmentation=true2022.08 | 61.1 | — | — | |
| LaSLIPArchitecture=ViT-B/16, Pre-training Data=CC12M, Protocol=5-way 5-shot2023.05 | 61 | — | — | |
| Supervised(IN1K)Pretrain Dataset=1.2M, Backbone=ResNet-50, Evaluation Protocol=End-to-end fine-tuning2022.04 | 60 | — | — | |
| CLIP*Image Encoder=ViT-B/16, Pretrain Dataset Size=400M, Evaluation Protocol=Linear probe2022.04 | 59.5 | — | — | |
| RényiCLEvaluation Protocol=linear evaluation, Multi-crop Augmentation=false2022.08 | 58.8 | — | — | |
| LaCLIPArchitecture=ViT-B/16, Pre-training Data=RedCaps, Protocol=5-way 5-shot2023.05 | 58.8 | — | — | |
| PyramidCLIPPretrain Dataset=143M, Backbone=ResNet-50, Evaluation Protocol=End-to-end fine-tuning2022.04 | 58.4 | — | — | |
| GalLoPShots=16, Backbone=ViT-B/16, Pre-training=CLIP2026.03 | 58.3 | — | — | |
| CLIP*Pretrain Dataset=400M, Backbone=ResNet-50, Evaluation Protocol=End-to-end fine-tuning2022.04 | 57.8 | — | — | |
| SOT-GLPShots=16, Backbone=ViT-B/16, Pre-training=CLIP2026.03 | 57.6 | — | — | |
| AWTShots=162024.07 | 57.08 | — | — | |
| PyramidCLIPImage Encoder=ViT-B/16, Pretrain Dataset Size=143M, Evaluation Protocol=Linear probe2022.04 | 56.5 | — | — | |
| SLIPArchitecture=ViT-B/16, Pre-training Data=CC12M, Protocol=5-way 5-shot2023.05 | 56.4 | — | — | |
| Skip TuningShot=16, Time (s)=962, Mem. (M)=5282024.12 | 55.57 | — | — | |
| Partial-LoRAParameter Count=0.31M2025.12 | 54.58 | — | — | |
| CLIPArchitecture=ViT-B/16, Pre-training Data=RedCaps, Protocol=5-way 5-shot2023.05 | 54.5 | — | — | |
| LoRAParameter Count=0.67M2025.12 | 54.34 | — | — | |
| DoRAParameter Count=0.75M2025.12 | 53.41 | — | — | |
| AdaLoRAParameter Count=0.67M2025.12 | 52.74 | — | — | |
| VeRAParameter Count=0.10M2025.12 | 52.29 | — | — | |
| Partial-AdaLoRAParameter Count=0.30M2025.12 | 52.22 | — | — | |
| CLIP*Image Encoder=ViT-B/32, Pretrain Dataset Size=400M, Evaluation Protocol=Linear probe2022.04 | 52 | — | — | |
| CLSAEvaluation Protocol=linear evaluation, Multi-crop Augmentation=false2022.08 | 51.6 | — | — | |
| SupervisedBackbone=Swin-T, Evaluation Protocol=Linear probe2021.06 | 51.5 | — | — | |
| LoRA+Parameter Count=0.67M2025.12 | 51.46 | — | — | |
| ProSRCShot=16, Time (s)=14403, Mem. (M)=33732024.12 | 50.83 | — | — | |
| PromptSRCShots=16, Backbone=ViT-B/16, Pre-training=CLIP2026.03 | 50.8 | — | — | |
| InfoMinEvaluation Protocol=linear evaluation, Multi-crop Augmentation=false2022.08 | 50.5 | — | — | |
| PyramidCLIPImage Encoder=ViT-B/32, Pretrain Dataset Size=143M, Evaluation Protocol=Linear probe2022.04 | 50.2 | — | — | |
| CLSAEvaluation Protocol=linear evaluation, Multi-crop Augmentation=true2022.08 | 50 | — | — | |
| LaCLIPArchitecture=ViT-B/16, Pre-training Data=CC3M, Protocol=5-way 5-shot2023.05 | 49.2 | — | — | |
| CLIPBackbone=ResNet-50, Evaluation Protocol=Linear probe2021.06 | 49.1 | — | — | |
| SwAVEvaluation Protocol=Linear probe, Backbone=ResNet-50, Pre-training Dataset=ImageNet, Multi-crop Augmentation=2 x 160 + 4 x 962022.09 | 49.1 | — | — | |
| SupervisedBackbone=ResNet-50, Evaluation Protocol=Linear probe2021.06 | 48.5 | — | — | |
| SupervisedBackbone=ResNet-50, Evaluation Protocol=Linear probe, Reproduced=true2021.06 | 48.4 | — | — | |
| MaPLeShot=16, Time (s)=14070, Mem. (M)=30612024.12 | 48.4 | — | — | |
| MaPLeShots=16, Backbone=ViT-B/16, Pre-training=CLIP2026.03 | 48.4 | — | — | |
| DISSLEvaluation Protocol=Linear probe, Backbone=ResNet-50, Pre-training Dataset=ImageNet, Multi-crop Augmentation=2 x 160 + 4 x 962022.09 | 48.1 | — | — | |
| SCANShot=162025.05 | 47.97 | — | — | |
| PLOTShots=16, Backbone=ViT-B/16, Pre-training=CLIP2026.03 | 46.7 | — | — | |
| Skip TuningShot=8, Time (s)=508, Mem. (M)=5282024.12 | 46.5 | — | — | |
| CLIPArchitecture=ViT-B/16, Pre-training Data=CC3M, Protocol=5-way 5-shot2023.05 | 46.1 | — | — | |
| Linear ProbeShots=16, Backbone=ViT-B/16, Pre-training=CLIP2026.03 | 45.4 | — | — | |
| CoOpShots=162024.07 | 43.4 | — | — | |
| CoOpShot=16, Time (s)=3879, Mem. (M)=45512024.12 | 43.4 | — | — | |
| CoOpShots=16, Backbone=ViT-B/16, Pre-training=CLIP2026.03 | 43.4 | — | — | |
| ProSRCShot=8, Time (s)=7636, Mem. (M)=33732024.12 | 43.27 | — | — | |
| OpenCLIP VIT-H/14zero-shot=true2023.03 | 42.3 | — | — | |
| MaPLeShot=8, Time (s)=7060, Mem. (M)=30612024.12 | 42 | — | — | |
| SCANShot=82025.05 | 41.4 | — | — | |
| StatABackbone=ViT-L/14, Scenario=Separate, Batch Size=1282025.01 | 41.3 | — | — | |
| StatABackbone=ViT-L/14, Batch Size=128, Online Test-Time Adaptation=true, Scenario=Separate2025.01 | 41.3 | — | — | |
| LoCoOpShots=16, Backbone=ViT-B/16, Pre-training=CLIP2026.03 | 40.7 | — | — | |
| Align+MixShot=16, Backbone=ResNet-502026.03 | 40.5 | — | — | |
| GDAShot=16, Backbone=ResNet-502026.03 | 40.3 | — | — | |
| SimNLShot=162025.05 | 40.27 | — | — | |
| ProDAShots=16, Backbone=ViT-B/16, Pre-training=CLIP2026.03 | 40.2 | — | — | |
| Skip TuningShot=4, Time (s)=280, Mem. (M)=5282024.12 | 39.9 | — | — | |
| StatABackbone=ViT-L/14, Scenario=High, Batch Size=1282025.01 | 39.6 | — | — | |
| StatABackbone=ViT-L/14, Batch Size=128, Online Test-Time Adaptation=true, Scenario=High2025.01 | 39.6 | — | — | |
| CoOpShot=8, Time (s)=1988, Mem. (M)=45512024.12 | 39 | — | — | |
| StatABackbone=ViT-L/14, Scenario=Medium, Batch Size=1282025.01 | 38.3 | — | — | |
| StatABackbone=ViT-L/14, Batch Size=128, Online Test-Time Adaptation=true, Scenario=Medium2025.01 | 38.3 | — | — | |
| ProSRCShot=4, Time (s)=3841, Mem. (M)=33732024.12 | 37.47 | — | — | |
| CLIPFT + ABackbone=ViT-B/32, Evaluation Protocol=Zero-shot, Text Prompting Strategy=LLM attributes, Training Approach=Fine-tuned2024.01 | 36.41 | — | — | |
| TaskResShot=162025.05 | 36.3 | — | — |