Object Detection on COCO mini (val)
62.5APSwinV2-G (HTC++)
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| SwinV2-G (HTC++)train I(W) size=1536(32), test I(W) size=ms(48), multi-scale testing=true2021.11 | 62.5 | — | — | — | — | — | — | |
| FlorenceInference Mode=Fine-tuning2021.11 | 62 | — | — | — | — | — | — | |
| SwinV2-G (HTC++)train I(W) size=1536(32), test I(W) size=1100(48), multi-scale testing=false2021.11 | 61.9 | — | — | — | — | — | — | |
| SwinV2-G (HTC++)train I(W) size=1536(32), test I(W) size=1100(32), multi-scale testing=false2021.11 | 61.7 | — | — | — | — | — | — | |
| Soft TeacherInference Mode=Fine-tuning2021.11 | 60.7 | — | — | — | — | — | — | |
| SoftTeachertrain I(W) size=1280(12), test I(W) size=ms(12), multi-scale testing=true2021.11 | 60.7 | — | — | — | — | — | — | |
| DyHeadInference Mode=Fine-tuning2021.11 | 60.3 | — | — | — | — | — | — | |
| DyHeadtrain I(W) size=1200(-), test I(W) size=ms(-), multi-scale testing=true2021.11 | 60.3 | — | — | — | — | — | — | |
| SwinV2-L (HTC++)train I(W) size=1536(32), test I(W) size=ms(48), multi-scale testing=true2021.11 | 60.2 | — | — | — | — | — | — | |
| CBNettrain I(W) size=1400(7), test I(W) size=ms(7), multi-scale testing=true2021.11 | 59.6 | — | — | — | — | — | — | |
| CB-Swin-L (HTC)Pre-trained on=ImageNet-22K, Params=453 M, Epochs=12, multi-scale testing=true2021.07 | 59.6 | — | — | — | — | — | — | |
| CB-Swin-L (HTC)Pre-trained on=ImageNet-22K, Params=453 M, Epochs=122021.07 | 59.1 | — | — | — | — | — | — | |
| SwinV2-L (HTC++)train I(W) size=1536(32), test I(W) size=1100(48), multi-scale testing=false2021.11 | 58.9 | — | — | — | — | — | — | |
| CB-Swin-B (HTC)Pre-trained on=ImageNet-22K, Params=235 M, Epochs=20, multi-scale testing=true2021.07 | 58.9 | — | — | — | — | — | — | |
| SwinV2-L (HTC++)train I(W) size=1536(32), test I(W) size=1100(32), multi-scale testing=false2021.11 | 58.8 | — | — | — | — | — | — | |
| Focal-LDetection Method=DyHead, Params=229M, Multi-scale evaluation=true2021.07 | 58.7 | — | — | — | — | — | — | |
| Swin-LDetection Method=DyHead, Params=213M, Multi-scale evaluation=true2021.07 | 58.4 | — | — | — | — | — | — | |
| CB-Swin-B (HTC)Pre-trained on=ImageNet-22K, Params=235 M, Epochs=202021.07 | 58.4 | — | — | — | — | — | — | |
| Focal-LDetection Method=HTC++, Params=265M, Multi-scale evaluation=true2021.07 | 58.1 | — | — | — | — | — | — | |
| Swin-LDetection Method=HTC++, Params=284M, Multi-scale evaluation=true2021.07 | 58 | — | — | — | — | — | — | |
| SwinV1-Ltrain I(W) size=800(7), test I(W) size=ms(7), multi-scale testing=true2021.11 | 58 | — | — | — | — | — | — | |
| Swin-L (HTC++)Pre-trained on=ImageNet-22K, Params=284 M, Epochs=72, multi-scale testing=true2021.07 | 58 | — | — | — | — | — | — | |
| Swin-L (HTC++)Backbone=Swin-L, Framework=HTC++, Multi-scale testing=true, #param.=284M2021.03 | 58 | — | — | — | — | — | — | |
| Swin-LDetection Method=HTC++, Params=284M, FLOPs=1470G, Multi-scale evaluation=false2021.07 | 57.1 | — | — | — | — | — | — | |
| Swin-L (HTC++)Pre-trained on=ImageNet-22K, Params=284 M, Epochs=722021.07 | 57.1 | — | — | — | — | — | — | |
| Swin-L (HTC++)Backbone=Swin-L, Framework=HTC++, #param.=284M, FLOPs=1470G2021.03 | 57.1 | — | — | — | — | — | — | |
| Focal-LDetection Method=HTC++, Params=265M, FLOPs=1165G, Multi-scale evaluation=false2021.07 | 57 | — | — | — | — | — | — | |
| CopyPastetrain I(W) size=1280(-), test I(W) size=1280(-), multi-scale testing=false2021.11 | 57 | — | — | — | — | — | — | |
| Focal-LDetection Method=DyHead, Params=229M, FLOPs=1081G, Multi-scale evaluation=false2021.07 | 56.4 | — | — | — | — | — | — | |
| Swin-B (HTC++)Backbone=Swin-B, Framework=HTC++, #param.=160M, FLOPs=1043G2021.03 | 56.4 | — | — | — | — | — | — | |
| CB-Swin-S (Cascade Mask R-CNN)Pre-trained on=ImageNet-1K, Params=156 M, Epochs=362021.07 | 56.3 | — | — | — | — | — | — | |
| Swin-LDetection Method=DyHead, Params=213M, FLOPs=965G, Multi-scale evaluation=false2021.07 | 56.2 | — | — | — | — | — | — | |
| Swin-LDetection Method=QueryInst, Multi-scale evaluation=true2021.07 | 56.1 | — | — | — | — | — | — | |
| Copy-pasteParams=185M, FLOPs=1440G, Multi-scale evaluation=false2021.07 | 55.9 | — | — | — | — | — | — | |
| CopyPastePre-trained on=ImageNet-1K, Params=185 M, Epochs=96, additional unlabeled data=true2021.07 | 55.9 | — | — | — | — | — | — | |
| Copy-paste#param.=185M, FLOPs=1440G2021.03 | 55.9 | — | — | — | — | — | — | |
| EfficientNet-D7Params=77M, FLOPs=410G, Multi-scale evaluation=false2021.07 | 54.4 | — | — | — | — | — | — | |
| EfficientDet-D7#param.=77M, FLOPs=410G2021.03 | 54.4 | — | — | — | — | — | — | |
| PET-DINOPrompt Type=Text, Backbone=Swin-L, Training Data=O365†, Zero-shot=true2026.04 | 54 | — | — | — | — | — | — | |
| MM-Grounding-DINOPrompt Type=Text, Backbone=Swin-L, Training Data=O365V2, OpenImage, GoldG, Zero-shot=true2026.04 | 53 | — | — | — | — | — | — | |
| SpineNet-190Params=164M, FLOPs=1885G, Multi-scale evaluation=false2021.07 | 52.6 | — | — | — | — | — | — | |
| SpineNet-190#param.=164M, FLOPs=1885G2021.03 | 52.6 | — | — | — | — | — | — | |
| ResNeSt-200Multi-scale evaluation=false2021.07 | 52.5 | — | — | — | — | — | — | |
| ResNeSt-200Pre-trained on=ImageNet-1K, Epochs=36, multi-scale testing=true2021.07 | 52.5 | — | — | — | — | — | — | |
| ResNeSt-200Multi-scale testing=true2021.03 | 52.5 | — | — | — | — | — | — | |
| Grounding DINOPrompt Type=Text, Backbone=Swin-L, Training Data=O365, OpenImages, GoldG, Zero-shot=true2026.04 | 52.5 | — | — | — | — | — | — | |
| X101-64x4dParams=155M, FLOPs=1033G, Multi-scale evaluation=false2021.07 | 52.3 | — | — | — | — | — | — | |
| X101-64 (HTC++)Backbone=ResNeXt-101-64x4d, Framework=HTC++, #param.=155M, FLOPs=1033G2021.03 | 52.3 | — | — | — | — | — | — | |
| Swin-BPre-trained on=ImageNet-1K, Params=145 M, Epochs=362021.07 | 51.9 | — | — | — | — | — | — | |
| GCNet*FLOPs=1041G, Multi-scale evaluation=false2021.07 | 51.8 | — | — | — | — | — | — | |
| GCNetPre-trained on=ImageNet-1K, Epochs=36, multi-scale testing=true2021.07 | 51.8 | — | — | — | — | — | — | |
| GCNetMulti-scale testing=true, FLOPs=1041G2021.03 | 51.8 | — | — | — | — | — | — | |
| MM-Grounding-DINOPrompt Type=Text, Backbone=Swin-T, Training Data=O365, GoldG, V3Det, Zero-shot=true2026.04 | 50.6 | — | — | — | — | — | — | |
| PET-DINOPrompt Type=Text, Backbone=Swin-T, Training Data=O365†, Zero-shot=true2026.04 | 49.8 | — | — | — | — | — | — | |
| BoTNet-200Multi-scale evaluation=false2021.07 | 49.7 | — | — | — | — | — | — | |
| PRISMambaFLOPs (G)=2352026.02 | 48.9 | 70.7 | 52.6 | — | — | — | — | |
| Swin-SBackbone=Swin-S, #params=69.1M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 48.5 | 70.2 | 53.5 | — | — | — | — | |
| XCiT-M24/8Backbone=XCiT-M24/8, #params=98.9M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 48.5 | 70.3 | 53.4 | — | — | — | — | |
| Grounding DINOPrompt Type=Text, Backbone=Swin-T, Training Data=O365, GoldG, Cap4M, Zero-shot=true2026.04 | 48.4 | — | — | — | — | — | — | |
| XCiT-S24/8Backbone=XCiT-S24/8, #params=64.5M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 48.1 | 69.5 | 53 | — | — | — | — | |
| GroupMambaFLOPs (G)=2792026.02 | 47.6 | 69.8 | 52.1 | — | — | — | — | |
| DefMambaFLOPs (G)=2682026.02 | 47.5 | 69.6 | 51.7 | — | — | — | — | |
| XCiT-S12/8Backbone=XCiT-S12/8, #params=43.1M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 47 | 68.9 | 51.7 | — | — | — | — | |
| VSSDFLOPs (G)=2652026.02 | 46.9 | 69.4 | 51.4 | — | — | — | — | |
| FractalMambaFLOPs (G)=2662026.02 | 46.8 | 68.7 | 50.8 | — | — | — | — | |
| XCiT-M24/16Backbone=XCiT-M24/16, #params=101.1M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 46.7 | 68.2 | 51.1 | — | — | — | — | |
| QuadMambaFLOPs (G)=3012026.02 | 46.7 | 69 | 51.3 | — | — | — | — | |
| LocalMambaFLOPs (G)=2912026.02 | 46.7 | 68.7 | 50.8 | — | — | — | — | |
| XCiT-S24/16Backbone=XCiT-S24/16, #params=65.8M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 46.5 | 68 | 50.9 | — | — | — | — | |
| VMambaFLOPs (G)=2622026.02 | 46.5 | 68.5 | 50.7 | — | — | — | — | |
| Adventurer2026.02 | 46.5 | 65.2 | 50.4 | — | — | — | — | |
| Swin-TBackbone=Swin-T, #params=47.8M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 46 | 68.1 | 50.3 | — | — | — | — | |
| PlainMamba-AdapterFLOPs (G)=5422026.02 | 46 | 66.9 | 50.1 | — | — | — | — | |
| ViL-LargeBackbone=ViL-Large, #params=76.1M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 45.7 | 67.2 | 49.9 | — | — | — | — | |
| Vim2026.02 | 45.7 | 63.9 | 49.6 | — | — | — | — | |
| XCiT-S12/16Backbone=XCiT-S12/16, #params=44.3M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 45.3 | 67 | 49.5 | — | — | — | — | |
| x^3 + dynamic scalePre-trained=ImageNet-1k, Model=Mask R-CNN2024.10 | 45.1 | 66.5 | 49.4 | — | — | — | — | |
| softmaxPre-trained=ImageNet-1k, Model=Mask R-CNN2024.10 | 44.9 | 66.1 | 48.9 | — | — | — | — | |
| x^3 + fixed scalePre-trained=ImageNet-1k, Model=Mask R-CNN2024.10 | 44.8 | 66.3 | 49.1 | — | — | — | — | |
| ViL-MediumBackbone=ViL-Medium, #params=60.1M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 44.6 | 66.3 | 48.5 | — | — | — | — | |
| PVT-LargeBackbone=PVT-Large, #params=81.0M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 44.5 | 66 | 48.3 | — | — | — | — | |
| ResNeXt101-64Backbone=ResNeXt101-64, #params=101.9M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 44.4 | 64.9 | 48.8 | — | — | — | — | |
| PVT-MediumBackbone=PVT-Medium, #params=63.9M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 44.2 | 66 | 48.2 | — | — | — | — | |
| ResNeXt101-32Backbone=ResNeXt101-32, #params=62.8M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 44 | 64.4 | 48 | — | — | — | — | |
| ViL-SmallBackbone=ViL-Small, #params=45.0M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 43.4 | 64.9 | 47 | — | — | — | — | |
| RevBiFPN-S6Backbone=RevBiFPN-S6, Params=130.2M, MACs=518.5B, Memory=7.4GB, Learning Schedule=1x, Framework=Mask R-CNN, Resolution=800x13332022.06 | 43.3 | — | — | 26.9 | 47.4 | 55.6 | — | |
| PVT-SmallBackbone=PVT-Small, #params=44.1M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 43 | 65.3 | 46.9 | — | — | — | — | |
| ResNet101Backbone=ResNet101, #params=63.2M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 42.8 | 63.2 | 47.1 | — | — | — | — | |
| HRNETV2P-W32Backbone=HRNETV2P-W32, Params=49.9M, MACs=352.0B, Memory=9.0GB, Learning Schedule=2x, Framework=Mask R-CNN, Resolution=800x13332022.06 | 42.3 | — | — | 25 | 45.4 | 54.9 | — | |
| RevBiFPN-S6Backbone=RevBiFPN-S6, Params=127.51M, MACs=465.43B, Mem=7.37 GB, LS=1x, Framework=Faster R-CNN2022.06 | 42.2 | 63.5 | 45.8 | 25.7 | 46.5 | 54 | — | |
| RevBiFPN-S5Backbone=RevBiFPN-S5, Params=80.5M, MACs=382.0B, Memory=5.5GB, Learning Schedule=1x, Framework=Mask R-CNN, Resolution=800x13332022.06 | 42.2 | — | — | 25.5 | 46.3 | 54.3 | — | |
| HRNETV2P-W48Backbone=HRNETV2P-W48, Params=83.36M, MACs=481.92B, Mem=11.64 GB, LS=2x, Framework=Faster R-CNN2022.06 | 41.8 | 62.8 | 45.9 | 25 | 44.7 | 54.6 | — | |
| XCiT-T12/8Backbone=XCiT-T12/8, #params=25.8M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 41.6 | 64.7 | 45.3 | — | — | — | — | |
| EfficientVMamba2026.02 | 41.6 | 63.2 | 45.3 | — | — | — | — | |
| RevBiFPN-S4Backbone=RevBiFPN-S4, Params=55.5M, MACs=304.1B, Memory=4.1GB, Learning Schedule=1x, Framework=Mask R-CNN, Resolution=800x13332022.06 | 41.5 | — | — | 24.2 | 45.4 | 53.9 | — | |
| RevBiFPN-S5Backbone=RevBiFPN-S5, Params=77.83M, MACs=328.91B, Mem=5.50 GB, LS=1x, Framework=Faster R-CNN2022.06 | 41.3 | 62.7 | 44.8 | 24.8 | 45.6 | 52.5 | — | |
| HRNETV2P-W48Backbone=HRNETV2P-W48, Params=83.36M, MACs=481.92B, Mem=11.64 GB, LS=1x, Framework=Faster R-CNN2022.06 | 41.3 | 62.8 | 45.1 | 25.1 | 44.5 | 52.9 | — | |
| ResNet50Backbone=ResNet50, #params=44.2M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 41 | 61.7 | 44.9 | — | — | — | — | |
| RESNET-101-FPNBackbone=RESNET-101-FPN, Params=63.2M, MACs=349.7B, Memory=5.8GB, Learning Schedule=2x, Framework=Mask R-CNN, Resolution=800x13332022.06 | 41 | — | — | 23.4 | 44.4 | 53.9 | — | |
| HRNETV2P-W32Backbone=HRNETV2P-W32, Params=47.28M, MACs=298.96B, Mem=8.62 GB, LS=2x, Framework=Faster R-CNN2022.06 | 40.9 | 61.8 | 44.8 | 24.4 | 43.7 | 53.3 | — |