Multi-label image recognition on MS-COCO 2014 (val)
91.3mAPML-Decoder + AAM
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| ML-Decoder + AAMBackbone=TResNet-L, Input resolution=640x640, GFLOPS=73.422022.09 | 91.3 | — | — | — | — | — | — | — | |
| ML-DecoderBackbone=TResNet-L, Input resolution=640x640, GFLOPS=73.422022.09 | 91.1 | — | — | — | — | — | — | — | |
| Q2LBackbone=TResNet-L, Input resolution=640x640, GFLOPS=119.692022.09 | 90.3 | — | — | — | — | — | — | — | |
| ML-Decoder + AAMBackbone=TResNet-L, Input resolution=448x448, GFLOPS=36.152022.09 | 90.3 | — | — | — | — | — | — | — | |
| ML-Decoder + AAMBackbone=EfficientNet-V2-L, Input resolution=448x448, GFLOPS=49.922022.09 | 90.1 | — | — | — | — | — | — | — | |
| ML-DecoderBackbone=TResNet-L, Input Resolution=448x4482021.11 | 90 | — | — | — | — | — | — | — | |
| ML-DecoderBackbone=TResNet-L, Input resolution=448x448, GFLOPS=36.152022.09 | 90 | — | — | — | — | — | — | — | |
| GAT re-weightingBackbone=TResNet-L, Input resolution=448x448, GFLOPS=35.22022.09 | 89.95 | — | — | — | — | — | — | — | |
| GATNBackbone=ResNeXt-101, Input resolution=448x448, GFLOPS=362022.09 | 89.3 | — | — | — | — | — | — | — | |
| Q2LBackbone=TResNet-L, Input Resolution=448x4482021.11 | 89.2 | — | — | — | — | — | — | — | |
| Q2LBackbone=TResNet-L, Input resolution=448x448, GFLOPS=60.42022.09 | 89.2 | — | — | — | — | — | — | — | |
| ML-Decoder + AAMBackbone=EfficientNet-V2-s, Input resolution=448x448, GFLOPS=12.282022.09 | 88.75 | — | — | — | — | — | — | — | |
| ASLBackbone=TResNet-L, Input Resolution=448x4482021.11 | 88.4 | — | — | — | — | — | — | — | |
| ASLBackbone=TResNet-L, Input resolution=448x448, GFLOPS=43.52022.09 | 88.4 | — | — | — | — | — | — | — | |
| ML-DecoderBackbone=EfficientNet-V2-s, Input resolution=448x448, GFLOPS=12.28, training_strategy=this_paper_training_strategy2022.09 | 88.25 | — | — | — | — | — | — | — | |
| GAT re-weightingBackbone=EfficientNet-V2-s, Input resolution=448x448, GFLOPS=10.832022.09 | 87.7 | — | — | — | — | — | — | — | |
| ML-GCNBackbone=EfficientNet-V2-s, Input resolution=448x448, GFLOPS=10.83, training_strategy=this_paper_training_strategy2022.09 | 87.5 | — | — | — | — | — | — | — | |
| Q2LBackbone=EfficientNet-V2-s, Input resolution=448x448, GFLOPS=16.25, training_strategy=this_paper_training_strategy2022.09 | 87.35 | — | — | — | — | — | — | — | |
| ML-DecoderBackbone=ResNet101, Input Resolution=448x4482021.11 | 87.1 | — | — | — | — | — | — | — | |
| ASLBackbone=EfficientNet-V2-s, Input resolution=448x448, GFLOPS=10.83, training_strategy=this_paper_training_strategy2022.09 | 87.05 | — | — | — | — | — | — | — | |
| ASLBackbone=TResNet-L, Input resolution=4482020.09 | 86.6 | — | — | 81.4 | — | — | 81.8 | — | |
| TResNet-L2020.03 | 86.4 | 87.6 | 76 | 81.4 | 88.4 | 78.9 | 83.4 | — | |
| ASLBackbone=ResNet101, Input resolution=4482020.09 | 85 | — | — | 80.3 | — | — | 82.3 | — | |
| ASLBackbone=ResNet101, Input Resolution=448x4482021.11 | 85 | — | — | — | — | — | — | — | |
| MS-CMAInput resolution=4482020.09 | 83.8 | — | — | 78.4 | — | — | 81 | — | |
| MCARInput resolution=4482020.09 | 83.8 | — | — | 78 | — | — | 80.3 | — | |
| SSGRLBackbone=ResNet101, Input Resolution=576x5762021.11 | 83.8 | — | — | — | — | — | — | — | |
| MS-CMABackbone=ResNet101, Input Resolution=448x4482021.11 | 83.8 | — | — | — | — | — | — | — | |
| KSSNetBackbone=ResNet101, GCN layers=4, Pre-training=ImageNet2019.11 | 83.7 | 84.6 | 73.2 | 77.2 | 87.8 | 76.2 | 81.5 | — | |
| KSSNetBackbone=ResNet1012020.03 | 83.7 | 84.6 | 73.2 | 77.2 | 87.8 | 76.2 | 81.5 | — | |
| KSSNetInput resolution=4482020.09 | 83.7 | — | — | 77.2 | — | — | 81.5 | — | |
| KSSNETBackbone=ResNet101, Input Resolution=448x4482021.11 | 83.7 | — | — | — | — | — | — | — | |
| Human Annotator2026.02 | 83.26 | — | — | — | — | — | — | — | |
| ML-GCNInput resolution=4482020.09 | 83 | — | — | 78 | — | — | 80.3 | — | |
| ML-GCNBackbone=ResNet101, Input Resolution=448x4482021.11 | 83 | — | — | — | — | — | — | — | |
| TagLLMLLM Backbone=Qwen3-VL-30B-A3B-Instruct2026.02 | 82.76 | 87.15 | 87.71 | 87.43 | 87.84 | 88.01 | 87.92 | 66 | |
| TagLLMLLM Backbone=Qwen3-VL-8B-Instruct2026.02 | 82.66 | 85.25 | 87.34 | 86.28 | 83.91 | 88.76 | 86.27 | 16 | |
| ML-GCN2019.11 | 82.4 | 84.4 | 71.4 | 77.4 | 85.8 | 74.5 | 79.8 | — | |
| CADMInput resolution=4482020.09 | 82.3 | — | — | 77 | — | — | 79.6 | — | |
| BPLLM Backbone=Qwen3-VL-8B-Instruct2026.02 | 80.76 | 77.11 | 90.28 | 83.18 | 75.77 | 91.68 | 82.97 | 166 | |
| BPLLM Backbone=Qwen3-VL-30B-A3B-Instruct2026.02 | 80.52 | 79.86 | 90.37 | 84.79 | 77.13 | 90.75 | 83.39 | 600 | |
| MOPLLM Backbone=Qwen3-VL-30B-A3B-Instruct2026.02 | 79.7 | 71.47 | 93.96 | 81.18 | 71.16 | 94.27 | 81.1 | 40 | |
| RAM++2026.02 | 79.05 | 89.53 | 61.66 | 73.03 | 89.23 | 55.08 | 68.12 | — | |
| MOPLLM Backbone=Qwen3-VL-8B-Instruct2026.02 | 78.8 | 67.59 | 93.81 | 78.57 | 65.63 | 94.75 | 77.54 | 12 | |
| ResNet1012019.11 | 77.3 | 80.2 | 66.7 | 72.8 | 83.9 | 70.8 | 76.8 | — | |
| SRN2019.11 | 77.1 | 81.6 | 65.4 | 71.2 | 82.7 | 69.9 | 75.8 | — | |
| TagCLIP2026.02 | 73.61 | 67.98 | 65.29 | 66.61 | 69.06 | 68.74 | 68.9 | 4 | |
| CLIP2026.02 | 66.85 | 65 | 61.65 | 63.28 | 59.3 | 59.2 | 59.25 | — | |
| CNN-RNN2019.11 | 61.2 | — | — | — | — | — | — | — | |
| NXTP2026.02 | 57.38 | 57.89 | 50.43 | 53.91 | 62.78 | 55.94 | 59.16 | 4 | |
| CaSED2026.02 | 54.42 | 83.3 | 28.7 | 42.69 | 86.22 | 23.85 | 37.36 | 3 | |
| Multi-Evidence2019.11 | — | 80.4 | 70.2 | 74.9 | 85.2 | 72.5 | 78.4 | — |