Long-tailed Image Classification on Places-LT (val test)
50.1Overall AccuracyVL-LTR
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| VL-LTRBackbone=ViT-B/16, Linguistic data=true2024.10 | 50.1 | 54.2 | 48.5 | 42 | |
| GNM-PTBackbone=ViT-B/16, Visual-only=true, DRW=true2024.10 | 50.1 | 46.6 | 53.3 | 49.4 | |
| GNM-PTBackbone=ViT-B/16, Visual-only=true, DRW=false2024.10 | 50 | 48.6 | 52.1 | 47.9 | |
| LPTBackbone=ViT-B/16, Visual-only=true, reproduced=true2024.10 | 49.7 | 47.6 | 52.1 | 48.4 | |
| RACBackbone=ViT-B/16, Linguistic data=true2024.10 | 47.2 | 48.7 | 48.3 | 41.8 | |
| DecoderBackbone=ViT-B/16, Visual-only=true2024.10 | 46.8 | — | — | — | |
| SHIKEBackbone=ResNet152, Type=DNN-based2024.10 | 41.9 | 43.6 | 39.2 | 44.8 | |
| NCLBackbone=ResNet152, Type=DNN-based2024.10 | 41.8 | — | — | — | |
| GPaCoBackbone=ResNet152, Type=DNN-based2024.10 | 41.7 | 39.5 | 47.2 | 33 | |
| LiVTBackbone=ViT-B/16, Visual-only=true2024.10 | 40.8 | 48.1 | 40.6 | 27.5 | |
| CCSAMBackbone=ResNet152, SAM=true2024.10 | 40.6 | 41.2 | 42.1 | 36.4 | |
| RIDEBackbone=ResNet152, Type=DNN-based2024.10 | 40.4 | 44.4 | 40.6 | 33 | |
| MisLASBackbone=ResNet152, Type=DNN-based2024.10 | 40.4 | 39.6 | 43.3 | 36.1 | |
| GCLBackbone=ResNet152, Type=DNN-based2024.10 | 40.3 | 38.6 | 42.6 | 38.4 | |
| LWSBackbone=ResNet152, Type=DNN-based2024.10 | 37.6 | 40.6 | 39.1 | 28.6 |