Fine-grained visual classification on CUB-200-2011 (test)
0.918Top-1 AccMP-FGVC
Evaluation Results
| Method | Links | |
|---|---|---|
| MP-FGVCBackbone=ViT-B-162023.09 | 0.918 | |
| SIM-TransBackbone=ViT-B-16, reproduced_from_source=true2023.09 | 0.913 | |
| TransFGBackbone=ViT-B-16, reproduced_from_source=true2023.09 | 0.911 | |
| IELTBackbone=ViT-B-16, reproduced_from_source=true2023.09 | 0.909 | |
| CALBackbone=CNN2023.09 | 0.906 | |
| API-NetBackbone=CNN2023.09 | 0.9 | |
| PRISBackbone=CNN2023.09 | 0.9 | |
| RLRR*Backbone=ViT-B/16, Pre-training=ImageNet-21k, Params. (M)=0.47, AugReg=true2024.03 | 0.898 | |
| SSFBackbone=ViT-B/16, Pre-training=ImageNet-21k, Params. (M)=0.39, AugReg=true2024.03 | 0.895 | |
| RLRRBackbone=ViT-B/16, Pre-training=ImageNet-21k, Params. (M)=0.47, AugReg=false2024.03 | 0.893 | |
| ARC*Backbone=ViT-B/16, Pre-training=ImageNet-21k, Params. (M)=0.25, AugReg=true2024.03 | 0.893 | |
| FDLBackbone=CNN2023.09 | 0.891 | |
| MSHQPBackbone=CNN2023.09 | 0.89 | |
| MRDMNBackbone=CNN2023.09 | 0.888 | |
| VPT-DeepBackbone=ViT-B/16, Pre-training=ImageNet-21k, Params. (M)=0.85, AugReg=false2024.03 | 0.885 | |
| ARCBackbone=ViT-B/16, Pre-training=ImageNet-21k, Params. (M)=0.25, AugReg=false2024.03 | 0.885 | |
| BiasBackbone=ViT-B/16, Pre-training=ImageNet-21k, Params. (M)=0.28, AugReg=false2024.03 | 0.884 | |
| LoRABackbone=ViT-B/16, Pre-training=ImageNet-21k, Params. (M)=0.44, AugReg=false2024.03 | 0.883 | |
| Full fine-tuningBackbone=ViT-B/16, Pre-training=ImageNet-21k, Params. (M)=85.98, AugReg=false2024.03 | 0.873 | |
| AdapterBackbone=ViT-B/16, Pre-training=ImageNet-21k, Params. (M)=0.41, AugReg=false2024.03 | 0.871 | |
| VPT-ShallowBackbone=ViT-B/16, Pre-training=ImageNet-21k, Params. (M)=0.25, AugReg=false2024.03 | 0.867 | |
| Linear probingBackbone=ViT-B/16, Pre-training=ImageNet-21k, Params. (M)=0.18, AugReg=false2024.03 | 0.853 | |
| SPT-DeepBackbone=ViT-B, Pre-training Method=MoCo-v3, Pre-training Dataset=ImageNet-1K, Tuning Depth=Deep2024.02 | 0.8447 | |
| SPT-ShallowBackbone=ViT-B, Pre-training Method=MoCo-v3, Pre-training Dataset=ImageNet-1K, Tuning Depth=Shallow2024.02 | 0.835 | |
| SaSPAType=Ours, Backbone=ResNet-50, Diffusion Model=BLIP-diffusion2024.06 | 0.832 | |
| SaSPA w/o BLIP-diffusionType=Ours, Backbone=ResNet-50, Diffusion Model=Stable Diffusion v1.52024.06 | 0.83 | |
| GateVPTBackbone=ViT-B, Pre-training Method=MoCo-v3, Pre-training Dataset=ImageNet-1K2024.02 | 0.8286 | |
| Real GuidanceType=Generative, Backbone=ResNet-502024.06 | 0.828 | |
| VPT-DeepBackbone=ViT-B, Pre-training Method=MoCo-v3, Pre-training Dataset=ImageNet-1K, Tuning Depth=Deep2024.02 | 0.8267 | |
| CAL-AugType=Traditional, Backbone=ResNet-502024.06 | 0.825 | |
| CAL-Aug + CutMixType=Traditional, Backbone=ResNet-502024.06 | 0.824 | |
| ALIAType=Generative, Backbone=ResNet-502024.06 | 0.82 | |
| CutMixType=Traditional, Backbone=ResNet-502024.06 | 0.818 | |
| FullBackbone=ViT-B, Pre-training Method=MoCo-v3, Pre-training Dataset=ImageNet-1K, Tuning Depth=Full2024.02 | 0.8175 | |
| No AugType=Traditional, Backbone=ResNet-502024.06 | 0.815 | |
| RandAugType=Traditional, Backbone=ResNet-502024.06 | 0.815 | |
| RandAug + CutMixType=Traditional, Backbone=ResNet-502024.06 | 0.812 | |
| QuICBackbone=VGG16, Resolution=448x4482026.01 | 0.81 | |
| FullBackbone=ViT-B, Pre-training Method=MAE, Pre-training Dataset=ImageNet-1K, Tuning Depth=Full2024.02 | 0.8055 | |
| QuICBackbone=ResNet18, Resolution=448x4482026.01 | 0.805 | |
| QuICBackbone=GoogLeNet, Resolution=448x4482026.01 | 0.803 | |
| SPT-DeepBackbone=ViT-B, Pre-training Method=MAE, Pre-training Dataset=ImageNet-1K, Tuning Depth=Deep2024.02 | 0.8013 | |
| VPT-ShallowBackbone=ViT-B, Pre-training Method=MoCo-v3, Pre-training Dataset=ImageNet-1K, Tuning Depth=Shallow2024.02 | 0.7905 | |
| GAPBackbone=ResNet18, Resolution=448x4482026.01 | 0.788 | |
| GAPBackbone=GoogLeNet, Resolution=448x4482026.01 | 0.784 | |
| SE-BlockBackbone=ResNet18, Resolution=448x4482026.01 | 0.777 | |
| GAPBackbone=VGG16, Resolution=448x4482026.01 | 0.77 | |
| FCBackbone=GoogLeNet, Resolution=448x4482026.01 | 0.767 | |
| SE-BlockBackbone=VGG16, Resolution=448x4482026.01 | 0.766 | |
| PANDStudent=MobileNet-V22026.02 | 0.7652 | |
| PANDStudent=ResNet-182026.02 | 0.7609 | |
| SE-BlockBackbone=GoogLeNet, Resolution=448x4482026.01 | 0.756 | |
| VL2LiteStudent=ResNet-182026.02 | 0.7267 | |
| VL2LiteStudent=MobileNet-V22026.02 | 0.7219 | |
| SPT-ShallowBackbone=ViT-B, Pre-training Method=MAE, Pre-training Dataset=ImageNet-1K, Tuning Depth=Shallow2024.02 | 0.7115 | |
| KDStudent=ResNet-182026.02 | 0.7095 | |
| GateVPTBackbone=ViT-B, Pre-training Method=MAE, Pre-training Dataset=ImageNet-1K2024.02 | 0.7056 | |
| VPT-DeepBackbone=ViT-B, Pre-training Method=MAE, Pre-training Dataset=ImageNet-1K, Tuning Depth=Deep2024.02 | 0.6833 | |
| RKDStudent=ResNet-182026.02 | 0.6831 | |
| KDStudent=MobileNet-V22026.02 | 0.68 | |
| RKDStudent=MobileNet-V22026.02 | 0.6795 | |
| w/o KDStudent=MobileNet-V22026.02 | 0.6542 | |
| w/o KDStudent=ResNet-182026.02 | 0.6448 | |
| FCBackbone=ResNet18, Resolution=448x4482026.01 | 0.635 | |
| QuICBackbone=AlexNet, Resolution=448x4482026.01 | 0.622 | |
| FCBackbone=VGG16, Resolution=448x4482026.01 | 0.613 | |
| GAPBackbone=AlexNet, Resolution=448x4482026.01 | 0.584 | |
| SE-BlockBackbone=AlexNet, Resolution=448x4482026.01 | 0.579 | |
| FCBackbone=AlexNet, Resolution=448x4482026.01 | 0.524 | |
| VPT-ShallowBackbone=ViT-B, Pre-training Method=MAE, Pre-training Dataset=ImageNet-1K, Tuning Depth=Shallow2024.02 | 0.4215 |