Person Re-identification on MSMT17 (test)
94.9Rank-1 AccCityGuard
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| CityGuardAttack Type=Clean2026.02 | 94.9 | 90.1 | — | — | |
| OATAttack Type=Clean2026.02 | 91.2 | 82.3 | — | — | |
| SOLIDERBackbone=Swin-S, Pre-training data=Large-scale Person Images, Number of pre-training images=4.2M, Input size=384 x 1282024.09 | 90.8 | 76.9 | — | — | |
| CLIP-REID + DenoiseRepBackbone=ViT-base2024.06 | 90.6 | 76.3 | — | — | |
| CIONBackbone=Swin-S, Pre-training data=Large-scale Person Images, Number of pre-training images=3.9M, Input size=384 x 1282024.09 | 90.4 | 77 | — | — | |
| PCL-CLIP Lpcl+LidBackbone=ViT, Supervision=Fully supervised, Resolution=256x1282023.10 | 89.8 | 76.1 | 94.7 | 96 | |
| CIONBackbone=R50-IBN, Pre-training data=Large-scale Person Images, Number of pre-training images=3.9M, Fine-tuning protocol=MGN, Input size=384 x 1282024.09 | 89.8 | 74.3 | — | — | |
| CLIP-REIDBackbone=ViT-base2024.06 | 89.7 | 75.8 | — | — | |
| TransReID-SSL + DenoiseRepBackbone=ViT-base2024.06 | 89.62 | 75.35 | — | — | |
| TransReID-SSLBackbone=ViT-base2024.06 | 89.5 | 75 | — | — | |
| PCL-CLIP LpclBackbone=ViT, Supervision=Fully supervised, Resolution=256x1282023.10 | 89.2 | 73.8 | 94.7 | 95.8 | |
| CLIP-REIDBackbone=ViT, Supervision=Fully supervised, Resolution=256x1282023.10 | 88.7 | 73.4 | — | — | |
| CLIP-ReIDBackbone=ViT, Learning Paradigm=Supervised, Venue=AAAI23, Camera Information Usage=false2026.01 | 88.7 | 73.4 | — | — | |
| CIONBackbone=ViT-B, Pre-training data=Large-scale Person Images, Number of pre-training images=3.9M2024.09 | 88.5 | 72.6 | — | — | |
| ISRBackbone=R50-IBN, Pre-training data=Large-scale Person Images, Number of pre-training images=47.8M, Fine-tuning protocol=MGN, Input size=384 x 1282024.09 | 88.4 | 71.5 | — | — | |
| LDSRe-ranking=true2021.11 | 88.35 | 79.09 | — | — | |
| ProNet++MG (Multi-granularity features)=true, #Params=1.43x, Re-ranking (RK)=true2023.08 | 88.2 | 80 | — | — | |
| PASSBackbone=ViT-B, Pre-training data=Large-scale Person Images, Number of pre-training images=4.2M2024.09 | 88.2 | 71.8 | — | — | |
| PASS*Pre-training=4 million pedestrian images, Input image size=256x1282023.06 | 88.2 | 71.8 | — | — | |
| TMGFBackbone=ViT, Learning Paradigm=Supervised, Venue=WACV23, Camera Information Usage=true2026.01 | 88.2 | 70.3 | 94.1 | 95.4 | |
| CIONBackbone=Swin-T, Pre-training data=Large-scale Person Images, Number of pre-training images=3.9M, Input size=384 x 1282024.09 | 88.1 | 71.1 | — | — | |
| CIONBackbone=R50, Pre-training data=Large-scale Person Images, Number of pre-training images=3.9M, Fine-tuning protocol=MGN, Input size=384 x 1282024.09 | 87.8 | 70.1 | — | — | |
| ProNet#Params=1x, Re-ranking (RK)=true2023.08 | 86.9 | 77.6 | — | — | |
| CIONBackbone=ViT-S, Pre-training data=Large-scale Person Images, Number of pre-training images=3.9M2024.09 | 86.9 | 70.3 | — | — | |
| IRM (MTL)Training Protocol=Multi-task Learning, Input image size=256x1282023.06 | 86.9 | 72.4 | — | — | |
| Ours(R101)+MGNBackbone=ResNet-101, Pre-training=LUPerson, Base Architecture=MGN2020.12 | 86.6 | 68.8 | — | — | |
| LDSRe-ranking=false2021.11 | 86.54 | 67.21 | — | — | |
| PASSBackbone=ViT-S, Pre-training data=Large-scale Person Images, Number of pre-training images=4.2M2024.09 | 86.5 | 69.1 | — | — | |
| FEDAttack Type=Clean2026.02 | 86.3 | 79.3 | — | — | |
| TransReIDBackbone=ViT, Supervision=Fully supervised, Resolution=256x1282023.10 | 86.2 | 69.4 | — | — | |
| BoT + CAJRe-ranking=CA-Jaccard2023.11 | 86.2 | 74.1 | 90.5 | — | |
| IRM (STL)Training Protocol=Single-task Learning, Input image size=256x1282023.06 | 86.2 | 71.9 | — | — | |
| TransReIDBackbone=ViT, Learning Paradigm=Supervised, Venue=ICCV21, Camera Information Usage=false2026.01 | 86.2 | 69.4 | — | — | |
| LUP NLBackbone=R50, Pre-training data=Large-scale Person Images, Number of pre-training images=10.7M, Fine-tuning protocol=MGN, Input size=384 x 1282024.09 | 86 | 68 | — | — | |
| SOLIDERBackbone=Swin-T, Pre-training data=Large-scale Person Images, Number of pre-training images=4.2M, Input size=384 x 1282024.09 | 85.9 | 67.4 | — | — | |
| SOLIDERReference=CVPR23, Transformer-based=true2024.10 | 85.9 | 67.4 | — | — | |
| ViTC-UReIDBackbone=ViT, Learning Paradigm=Unsupervised, Venue=Our, Camera Information Usage=false2026.01 | 85.8 | 63.6 | 92.3 | 94.1 | |
| TransReID + DenoiseRepBackbone=ViT-base-ics2024.06 | 85.72 | 68.1 | — | — | |
| MGN(R101)Backbone=ResNet-1012020.12 | 85.7 | 66 | — | — | |
| GiT2021.07 | 85.6 | 64.8 | — | — | |
| Ours(R50)+MGNBackbone=ResNet-50, Pre-training=LUPerson, Base Architecture=MGN2020.12 | 85.5 | 65.7 | — | — | |
| AdaSPBackbone=ResNet-50, Input size=384×1282023.03 | 85.5 | 67.1 | — | — | |
| LUPBackbone=R50, Pre-training data=Large-scale Person Images, Number of pre-training images=4.2M, Fine-tuning protocol=MGN, Input size=384 x 1282024.09 | 85.5 | 65.7 | — | — | |
| TranSSLBackbone=ViT-S, Pre-training data=Large-scale Person Images, Number of pre-training images=4.2M2024.09 | 85.5 | 66.8 | — | — | |
| TransReID-SSL + DenoiseRepBackbone=ViT-small2024.06 | 85.5 | 67.33 | — | — | |
| ProNet++MG (Multi-granularity features)=true, #Params=1.43x2023.08 | 85.4 | 65.5 | — | — | |
| ProNet++Backbone=CNN, Learning Paradigm=Supervised, Venue=CoRR23, Camera Information Usage=false2026.01 | 85.4 | 65.5 | — | — | |
| PLIPBackbone=R50, Pre-training data=Large-scale Person Images, Number of pre-training images=4.8M, Fine-tuning protocol=MGN, Input size=384 x 1282024.09 | 85.3 | 66.2 | — | — | |
| TransReIDBackbone=ViT-base-ics2024.06 | 85.3 | 67.7 | — | — | |
| MGN*Backbone=ResNet-50, Framework=fast-reid2020.12 | 85.1 | 63.7 | — | — | |
| MGNBackbone=R50, Pre-training data=ImageNet1K-1.3M, Input size=384 x 1282024.09 | 85.1 | 63.7 | — | — | |
| PCL-CLIP O2CAPBackbone=ViT, Supervision=Unsupervised, Resolution=256x1282023.10 | 84.9 | 65.5 | 92 | 94 | |
| PCL-CLIPBackbone=ViT, Learning Paradigm=Unsupervised, Venue=Arxiv23, Camera Information Usage=true2026.01 | 84.9 | 65.5 | 92 | 94 | |
| ReMixReference=Ours, Single-camera data usage=true2024.10 | 84.8 | 63.9 | — | — | |
| TransReID-SSLBackbone=ViT-small2024.06 | 84.8 | 66.3 | — | — | |
| CLIP-ReIDReference=AAAI232024.10 | 84.4 | 63 | — | — | |
| CLIP-ReID (AAAI'22)Backbone=ResNet-50, Resolution=256x1282024.05 | 84.4 | 63 | — | — | |
| AdaSPBackbone=ResNet-50, Input size=256×1282023.03 | 84.3 | 64.7 | — | — | |
| UPReIDBackbone=R50, Pre-training data=Large-scale Person Images, Number of pre-training images=4.2M, Fine-tuning protocol=MGN, Input size=384 x 1282024.09 | 84.3 | 63.3 | — | — | |
| PATHBackbone=ViT-B, Pre-training data=Large-scale Person Images, Number of pre-training images=11M2024.09 | 84.3 | 69.1 | — | — | |
| AdaSPReference=CVPR232024.10 | 84.3 | 64.7 | — | — | |
| Baseline† + CALBackbone=ResNet-50, Status=Strong Baseline2021.08 | 84.2 | 64 | 92 | — | |
| DSFLPublication=ACM MM20, Re-ranking=false2021.11 | 84.2 | 60.7 | — | — | |
| CALBackbone=CNN, Supervision=Fully supervised, Resolution=384x1922023.10 | 84.2 | 64 | — | — | |
| ReMix (w/o s-cam.)Reference=Ours, Single-camera data usage=false2024.10 | 83.9 | 62.8 | — | — | |
| SCSNCategory=Attention2020.09 | 83.8 | 58.5 | — | — | |
| SCSN2021.07 | 83.8 | 58.5 | — | — | |
| CLIP-ReID Baseline + OursBackbone=ResNet-50, Resolution=256x1282024.05 | 83.8 | 67.6 | — | — | |
| AAformerInput size=256x1282022.05 | 83.6 | 63.2 | — | — | |
| AAformerBackbone=ViT, Supervision=Fully supervised, Resolution=384x2562023.10 | 83.6 | 63.2 | — | — | |
| CACE-NetCategory=CACE-Net2020.09 | 83.54 | 62 | — | — | |
| MPNMethod Category=Part-based2020.03 | 83.5 | 62.7 | — | — | |
| TransReIDAttack Type=Clean2026.02 | 83.5 | 64.4 | — | — | |
| TransReIDMG (Multi-granularity features)=true, #Params=4.14x, Extra data or annotations (*)=true, Architecture (Transformer)=true2023.08 | 83.3 | 64.9 | — | — | |
| TMGFWBackbone=ViT, Supervision=Unsupervised, Resolution=256x1282023.10 | 83.3 | 58.2 | 90.2 | 92.1 | |
| FlipReIDReference=EUVIP212024.10 | 83.3 | 64.3 | — | — | |
| FastReIDBackbone=ResNet-50, Resolution=256x1282024.05 | 83.3 | 59.9 | — | — | |
| TMGFBackbone=ViT, Learning Paradigm=Unsupervised, Venue=WACV23, Camera Information Usage=true2026.01 | 83.3 | 58.2 | 90.2 | 92.1 | |
| MGNSupervision=Fully2019.04 | 83.1 | — | — | — | |
| DeiT-Base + DCALBackbone=DeiT-Base, Input size=256x1282022.05 | 83.1 | 62.3 | — | — | |
| ViT-Base + DCALBackbone=ViT-Base, Input size=256x1282022.05 | 83.1 | 64 | — | — | |
| DCAL#Params=> 4.14x, Extra data or annotations (*)=true, Architecture (Transformer)=true2023.08 | 83.1 | 64 | — | — | |
| DCALBackbone=ViT-B, Pre-training data=ImageNet21K-14M2024.09 | 83.1 | 64 | — | — | |
| SCSN (3 stage)Publication=CVPR20, Re-ranking=false2021.11 | 83 | 58 | — | — | |
| DFF-ResNet-50Backbone=ResNet-502022.03 | 82.95 | 60.21 | — | — | |
| Local CNNSupervision=Fully2019.04 | 82.9 | — | — | — | |
| ProNet#Params=1x2023.08 | 82.9 | 61.3 | — | — | |
| TransReIDBackbone=ViT-Base, Side Information=false, Input size=256x1282022.05 | 82.5 | 63.6 | — | — | |
| TransReIDBackbone=ViT-B, Pre-training data=ImageNet21K-14M2024.09 | 82.5 | 63.6 | — | — | |
| ABD-Net#Params=2.76x2023.08 | 82.4 | 60.8 | — | — | |
| CIONBackbone=R50-IBN, Pre-training=Large-scale Person Images2024.09 | 82.4 | 59.4 | — | — | |
| CIONBackbone=R50-IBN, Pre-training=Large-scale Person Images2024.09 | 82.4 | 59.4 | — | — | |
| CIONBackbone=R50-IBN, Pre-training=Large-scale Person Images2024.09 | 82.4 | 59.4 | — | — | |
| ABD-Net2019.08 | 82.3 | 60.8 | 90.6 | — | |
| ABD-NetSupervision=Fully2019.04 | 82.3 | — | — | — | |
| ABD-NetCategory=Attention2020.09 | 82.3 | 60.8 | — | — | |
| ABD-NetSetting=Fully Supervised2020.12 | 82.3 | 60.8 | 90.6 | — | |
| ABDNetBackbone=ResNet-502020.12 | 82.3 | 60.8 | — | — | |
| ABD-NetPublication=ICCV19, Re-ranking=false2021.11 | 82.3 | 60.8 | — | — | |
| ABDNetInput size=256x1282022.05 | 82.3 | 60.8 | — | — |