Long-tail Image Classification on iNaturalist 2018 (test)
88.3Accuracy (Few)LIFT
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| LIFTBackbone=ViT-L/14@336px, Learnable Params.=6.37M, #Epochs=20, TTE=true2023.09 | 88.3 | 83.6 | 87.4 | — | 87.4 | |
| LIFTBackbone=ViT-L/14@336px, Learnable Params.=6.37M, #Epochs=20, TTE=false2023.09 | 87.9 | 83.2 | 87 | — | 87 | |
| Category ExtrapolationBackbone=ViT-B, Pre-training paradigm=Fine-tuning pre-trained model (DINOv2)2024.10 | 86.7 | 86.4 | 87.4 | — | 87 | |
| RACInput Resolution=384x3842022.02 | 86.06 | 82.91 | 85.71 | 85.56 | — | |
| Bal-CE†Backbone=ViT-B, Pre-training paradigm=Fine-tuning pre-trained model (DINOv2)2024.10 | 85 | 85.8 | 86.5 | — | 85.9 | |
| Bal-CEBackbone=ViT-B, Pre-training paradigm=Fine-tuning pre-trained model (DINOv2)2024.10 | 84.2 | 85.7 | 86.2 | — | 85 | |
| LIFTBackbone=ViT-B/16, Learnable Params.=4.75M, #Epochs=20, Protocol=Fine-tuning foundation model, Test-time ensembling (TTE)=true2023.09 | 82.2 | 74 | 80.3 | — | 80.4 | |
| Category ExtrapolationBackbone=ViT-B, Pre-training paradigm=Fine-tuning pre-trained model (CLIP)2024.10 | 82.1 | 79.6 | 80.1 | — | 80.9 | |
| LIFT†Backbone=ViT-B, Pre-training paradigm=Fine-tuning pre-trained model (CLIP)2024.10 | 81.3 | 72.9 | 79.4 | — | 79.5 | |
| LIFTBackbone=ViT-B/16, Learnable Params.=4.75M, #Epochs=20, Protocol=Fine-tuning foundation model, Test-time ensembling (TTE)=false2023.09 | 81.1 | 72.4 | 79 | — | 79.1 | |
| RACBackbone=ViT-B/16, Learnable Params.=85.80M, #Epochs=20, Protocol=Fine-tuning with extra data2023.09 | 81.1 | 75.9 | 80.5 | — | 80.2 | |
| LIFTBackbone=ViT-B, Pre-training paradigm=Fine-tuning pre-trained model (CLIP)2024.10 | 81.1 | 72.4 | 79 | — | 79.1 | |
| RACInput Resolution=224x2242022.02 | 81.07 | 75.92 | 80.47 | 80.24 | — | |
| LPTBackbone=ViT-B/16, Learnable Params.=1.01M, #Epochs=80+80, Protocol=Fine-tuning foundation model2023.09 | 79.3 | — | — | — | 76.1 | |
| Category ExtrapolationBackbone=ViT-B, Pre-training paradigm=Training from scratch2024.10 | 77.5 | 78.9 | 78.2 | — | 78 | |
| LiVT†Backbone=ViT-B, Pre-training paradigm=Training from scratch2024.10 | 75.9 | 78.8 | 77.4 | — | 77 | |
| LiVTBackbone=ViT-B/16, Learnable Params.=85.80M, #Epochs=100, Protocol=Training from scratch2023.09 | 74.8 | 78.9 | 76.5 | — | 76.1 | |
| LiVTBackbone=ViT-B, Pre-training paradigm=Training from scratch2024.10 | 74.8 | 78.9 | 76.5 | — | 76.1 | |
| PaCoInput Resolution=224x2242022.02 | 74.7 | 75 | 75.5 | 75.2 | — | |
| RIDE+CRBackbone=ResNet-50, Learnable Params.=23.51M, #Epochs=200, Protocol=Training from scratch2023.09 | 74.3 | 71 | 73.8 | — | 73.5 | |
| RIDE+OTmixBackbone=ResNet-50, Learnable Params.=23.51M, #Epochs=210, Protocol=Training from scratch2023.09 | 73.8 | 71.3 | 72.8 | — | 73 | |
| NCLBackbone=ResNet-50, Learnable Params.=23.51M, #Epochs=400, Protocol=Training from scratch2023.09 | 73.8 | 72 | 74.9 | — | 74.2 | |
| PaCoBackbone=ResNet-50, Learnable Params.=23.51M, #Epochs=400, Protocol=Training from scratch2023.09 | 73.6 | 70.4 | 72.8 | — | 73.2 | |
| RIDEInput Resolution=224x224, Number of experts=42022.02 | 73.1 | 70.9 | 72.4 | 72.6 | — | |
| TADEInput Resolution=224x2242022.02 | 73.1 | 74.4 | 72.5 | 72.9 | — | |
| PaCo2022.03 | 73.1 | 69.5 | 72.3 | 72.3 | — | |
| RIDEBackbone=ResNet-50, Learnable Params.=23.51M, #Epochs=100, Protocol=Training from scratch2023.09 | 73.1 | 70.9 | 72.4 | — | 72.6 | |
| RIDEInput Resolution=224x224, Number of experts=22022.02 | 71.7 | 70.2 | 71.3 | 71.4 | — | |
| RIDE2022.03 | 71.5 | 66.5 | 72.1 | 71.3 | — | |
| LCReglatent_category_features=true, reconstruction_loss=true, latent_augmentation_loss=true2022.06 | 71.5 | 73.8 | 73.4 | — | — | |
| ALAInput Resolution=224x2242022.02 | 70.4 | 71.3 | 70.8 | 70.7 | — | |
| LCReg (baseline)latent_category_features=false, reconstruction_loss=false, latent_augmentation_loss=false2022.06 | 70.4 | 73.2 | 72.4 | — | — | |
| MiSLASBackbone=ResNet-50, Learnable Params.=23.51M, #Epochs=200+30, Protocol=Training from scratch2023.09 | 70.4 | 73.2 | 72.4 | — | 71.6 | |
| ALABackbone=ResNet-50, Learnable Params.=23.51M, #Epochs=90, Protocol=Training from scratch2023.09 | 70.4 | 71.3 | 70.8 | — | 70.7 | |
| DisAlign2022.03 | 70.2 | 69 | 71.1 | 70.6 | — | |
| DisAlignBackbone=ResNet-50, Learnable Params.=23.51M, #Epochs=90, Protocol=Training from scratch2023.09 | 69.9 | 61.6 | 70.8 | — | 69.5 | |
| Weight Balancingvariant=WD + WD & Max2022.03 | 69.7 | 71.2 | 70.4 | 70.2 | — | |
| Weight Balancingvariant=WD + WD2022.03 | 69.4 | 71 | 70.3 | 70 | — | |
| T-norm2022.06 | 69.3 | 71.1 | 68.9 | — | — | |
| Weight Balancingvariant=WD + Max2022.03 | 69.1 | 71.4 | 68.9 | 69.2 | — | |
| Weight Balancingvariant=WD + τ-norm2022.03 | 68.9 | 71.3 | 69.8 | 69.6 | — | |
| LWS2022.06 | 68.8 | 71 | 69.8 | — | — | |
| DiVE2022.03 | 67.6 | 70.6 | 70 | 69.1 | — | |
| DiVEBackbone=ResNet-50, Learnable Params.=23.51M, #Epochs=90, Protocol=Training from scratch2023.09 | 67.6 | 70.6 | 70 | — | 69.1 | |
| Weight Balancingvariant=WD + L2norm2022.03 | 66.9 | 71.2 | 47.4 | 51.3 | — | |
| CRT2022.06 | 66.1 | 73.2 | 68.8 | — | — | |
| τ-norm2022.03 | 65.5 | 65.6 | 65.3 | 65.6 | — | |
| LWSBackbone=ResNet-50, Learnable Params.=23.51M, #Epochs=90+10, Protocol=Training from scratch2023.09 | 65.5 | 65 | 66.3 | — | 65.9 | |
| BBN2022.03 | 65.3 | 49.4 | 70.8 | 66.3 | — | |
| OLTRInput Resolution=224x2242022.02 | 64.9 | 59 | 64.1 | 63.9 | — | |
| OLTR2022.03 | 64.9 | 59 | 64.1 | 63.9 | — | |
| cRT2022.03 | 63.2 | 69 | 66 | 65.2 | — | |
| cRTBackbone=ResNet-50, Learnable Params.=23.51M, #Epochs=90+10, Protocol=Training from scratch2023.09 | 63.2 | 69 | 66 | — | 65.2 | |
| Weight Balancingvariant=WD2022.03 | 61.5 | 74.5 | 66.5 | 65.4 | — | |
| KD2022.03 | 57.4 | 72.6 | 63.8 | 62.2 | — | |
| CE2022.03 | 57.2 | 72.2 | 63 | 61.7 | — | |
| CE+CB2022.03 | 53.2 | 53.4 | 54.8 | 54 | — | |
| Weight Balancingvariant=naive2022.03 | 43.9 | 54.7 | 46 | 46.1 | — | |
| Zero-shot CLIPBackbone=ViT-L/14@336px2023.09 | 5.3 | 11.5 | 5.8 | — | 6.2 | |
| ABC Norm2024.07 | — | — | — | — | 71.4 | |
| ACE2022.03 | — | — | — | 72.9 | — | |
| ALA2024.07 | — | — | — | — | 70.7 | |
| APA*LayerNorm=True, Attention dropout=0.12024.07 | — | — | — | — | 72.3 | |
| APA* + AGLULayerNorm=True, Attention dropout=0.1, Activation=AGLU2024.07 | — | — | — | — | 74.8 | |
| AREA2024.07 | — | — | — | — | 68.4 | |
| Baseline with SE2024.07 | — | — | — | — | 71.3 | |
| BCLBackbone=ResNet-50, Learnable Params.=23.51M, #Epochs=100, Protocol=Training from scratch2023.09 | — | — | — | — | 71.8 | |
| BCL2024.07 | — | — | — | — | 71.8 | |
| CC-SAM2024.07 | — | — | — | — | 70.9 | |
| DecoderBackbone=ViT-B/16, Learnable Params.=21.26M, #Epochs=~5, Protocol=Fine-tuning foundation model2023.09 | — | — | — | — | 59.2 | |
| Decouple-LWSInput Resolution=224x2242022.02 | — | — | — | 65.9 | — | |
| DisAlignInput Resolution=224x2242022.02 | — | — | — | 74.1 | — | |
| DisAlign2024.07 | — | — | — | — | 70.6 | |
| DOC2024.07 | — | — | — | — | 71 | |
| DRO-LT2022.03 | — | — | — | 69.7 | — | |
| Focal2022.03 | — | — | — | 61.1 | — | |
| GMLBackbone=ResNet-50, Learnable Params.=23.51M, #Epochs=400, Protocol=Training from scratch2023.09 | — | — | — | — | 74.5 | |
| GML lossBackbone=ViT-B2023.05 | — | — | — | — | 82.1 | |
| GrafitInput Resolution=224x2242022.02 | — | — | — | 69.9 | — | |
| GrafitInput Resolution=384x3842022.02 | — | — | — | 81.2 | — | |
| LACEInput Resolution=224x2242022.02 | — | — | — | 71.9 | — | |
| LADEInput Resolution=224x2242022.02 | — | — | — | 69.3 | — | |
| LADE2024.07 | — | — | — | — | 70 | |
| LWS+ImbSAM2024.07 | — | — | — | — | 71.1 | |
| MisLAS2024.07 | — | — | — | — | 71.6 | |
| ResLT2024.07 | — | — | — | — | 70.5 | |
| SSD2022.03 | — | — | — | 71.5 | — | |
| TSC2024.07 | — | — | — | — | 69.7 | |
| VL-LTRBackbone=ViT-B2023.05 | — | — | — | — | 81 | |
| VL-LTRBackbone=ViT-B/16, Learnable Params.=149.62M, #Epochs=100, Protocol=Fine-tuning with extra data2023.09 | — | — | — | — | 76.8 | |
| WD+MaxNorm2024.07 | — | — | — | — | 70.2 |