Instance Segmentation on COCO mini (val)
53.7AP^mSwinV2-G (HTC++)
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| SwinV2-G (HTC++)train I(W) size=1536(32), test I(W) size=ms(48), multi-scale testing=true2021.11 | 53.7 | — | — | — | — | — | |
| SwinV2-G (HTC++)train I(W) size=1536(32), test I(W) size=1100(48), multi-scale testing=false2021.11 | 53.4 | — | — | — | — | — | |
| SwinV2-G (HTC++)train I(W) size=1536(32), test I(W) size=1100(32), multi-scale testing=false2021.11 | 53.3 | — | — | — | — | — | |
| RevCol-H (HTC++)Backbone=RevCol-H, Pre-training=Extra data, Detection Framework=HTC++, Params=2.41G, FLOPs=4417G, Input Resolution=1280x800, Testing Scale=Single-scale2022.12 | 53 | 76.3 | 58.7 | — | — | — | |
| SoftTeachertrain I(W) size=1280(12), test I(W) size=ms(12), multi-scale testing=true2021.11 | 52.5 | — | — | — | — | — | |
| SwinV2-L (HTC++)train I(W) size=1536(32), test I(W) size=ms(48), multi-scale testing=true2021.11 | 52.1 | — | — | — | — | — | |
| CBNettrain I(W) size=1400(7), test I(W) size=ms(7), multi-scale testing=true2021.11 | 51.8 | — | — | — | — | — | |
| CB-Swin-L (HTC)Pre-trained on=ImageNet-22K, Params=453 M, Epochs=12, multi-scale testing=true2021.07 | 51.8 | — | — | — | — | — | |
| CB-Swin-B (HTC)Pre-trained on=ImageNet-22K, Params=235 M, Epochs=20, multi-scale testing=true2021.07 | 51.3 | — | — | — | — | — | |
| SwinV2-L (HTC++)train I(W) size=1536(32), test I(W) size=1100(48), multi-scale testing=false2021.11 | 51.2 | — | — | — | — | — | |
| SwinV2-L (HTC++)train I(W) size=1536(32), test I(W) size=1100(32), multi-scale testing=false2021.11 | 51.1 | — | — | — | — | — | |
| CB-Swin-L (HTC)Pre-trained on=ImageNet-22K, Params=453 M, Epochs=122021.07 | 51 | — | — | — | — | — | |
| Focal-LDetection Method=HTC++, Params=265M, Multi-scale evaluation=true2021.07 | 50.9 | — | — | — | — | — | |
| CB-Swin-B (HTC)Pre-trained on=ImageNet-22K, Params=235 M, Epochs=202021.07 | 50.7 | — | — | — | — | — | |
| Swin-LDetection Method=HTC++, Params=284M, Multi-scale evaluation=true2021.07 | 50.4 | — | — | — | — | — | |
| SwinV1-Ltrain I(W) size=800(7), test I(W) size=ms(7), multi-scale testing=true2021.11 | 50.4 | — | — | — | — | — | |
| Swin-L (HTC++)Pre-trained on=ImageNet-22K, Params=284 M, Epochs=72, multi-scale testing=true2021.07 | 50.4 | — | — | — | — | — | |
| Focal-LDetection Method=HTC++, Params=265M, FLOPs=1165G, Multi-scale evaluation=false2021.07 | 49.9 | — | — | — | — | — | |
| Swin-LDetection Method=HTC++, Params=284M, FLOPs=1470G, Multi-scale evaluation=false2021.07 | 49.5 | — | — | — | — | — | |
| Swin-L (HTC++)Pre-trained on=ImageNet-22K, Params=284 M, Epochs=722021.07 | 49.5 | — | — | — | — | — | |
| CopyPastetrain I(W) size=1280(-), test I(W) size=1280(-), multi-scale testing=false2021.11 | 48.9 | — | — | — | — | — | |
| CB-Swin-S (Cascade Mask R-CNN)Pre-trained on=ImageNet-1K, Params=156 M, Epochs=362021.07 | 48.6 | — | — | — | — | — | |
| RevCol-LBackbone=RevCol-L, Pre-training=ImageNet-22K, Params=330M, FLOPs=1453G, Input Resolution=1280x800, Testing Scale=Single-scale2022.12 | 48.4 | 71.8 | 52.8 | — | — | — | |
| ConvNeXt-LBackbone=ConvNeXt-L, Pre-training=ImageNet-22K, Params=255M, FLOPs=1354G, Input Resolution=1280x800, Testing Scale=Single-scale2022.12 | 47.6 | 71.3 | 51.7 | — | — | — | |
| RevCol-BBackbone=RevCol-B, Pre-training=ImageNet-22K, Params=196M, FLOPs=988G, Input Resolution=1280x800, Testing Scale=Single-scale2022.12 | 47.5 | 71.1 | 51.8 | — | — | — | |
| Copy-pasteParams=185M, FLOPs=1440G, Multi-scale evaluation=false2021.07 | 47.2 | — | — | — | — | — | |
| CopyPastePre-trained on=ImageNet-1K, Params=185 M, Epochs=96, additional unlabeled data=true2021.07 | 47.2 | — | — | — | — | — | |
| ConvNeXt-BBackbone=ConvNeXt-B, Pre-training=ImageNet-22K, Params=146M, FLOPs=964G, Input Resolution=1280x800, Testing Scale=Single-scale2022.12 | 46.9 | 70.6 | 51.3 | — | — | — | |
| Swin-LBackbone=Swin-L, Pre-training=ImageNet-22K, Params=253M, FLOPs=1382G, Input Resolution=1280x800, Testing Scale=Single-scale2022.12 | 46.7 | 70.1 | 50.8 | — | — | — | |
| RepLKNet-LBackbone=RepLKNet-L, Pre-training=ImageNet-22K, Params=229M, FLOPs=1321G, Input Resolution=1280x800, Testing Scale=Single-scale2022.12 | 46.5 | — | — | — | — | — | |
| RepLKNet-BBackbone=RepLKNet-B, Pre-training=ImageNet-22K, Params=137M, FLOPs=965G, Input Resolution=1280x800, Testing Scale=Single-scale2022.12 | 46.3 | — | — | — | — | — | |
| X101-64x4dParams=155M, FLOPs=1033G, Multi-scale evaluation=false2021.07 | 46 | — | — | — | — | — | |
| RevCol-BBackbone=RevCol-B, Pre-training=ImageNet-1K, Params=196M, FLOPs=988G, Input Resolution=1280x800, Testing Scale=Single-scale2022.12 | 45.9 | 69.1 | 50.1 | — | — | — | |
| Swin-BBackbone=Swin-B, Pre-training=ImageNet-22K, Params=145M, FLOPs=982G, Input Resolution=1280x800, Testing Scale=Single-scale2022.12 | 45.8 | 69.4 | 49.7 | — | — | — | |
| ConvNeXt-BBackbone=ConvNeXt-B, Pre-training=ImageNet-1K, Params=146M, FLOPs=964G, Input Resolution=1280x800, Testing Scale=Single-scale2022.12 | 45.6 | 68.9 | 49.5 | — | — | — | |
| RevCol-SBackbone=RevCol-S, Pre-training=ImageNet-1K, Params=118M, FLOPs=833G, Input Resolution=1280x800, Testing Scale=Single-scale2022.12 | 45.5 | 68.8 | 49 | — | — | — | |
| RepLKNet-BBackbone=RepLKNet-B, Pre-training=ImageNet-1K, Params=137M, FLOPs=965G, Input Resolution=1280x800, Testing Scale=Single-scale2022.12 | 45.2 | — | — | — | — | — | |
| Swin-BPre-trained on=ImageNet-1K, Params=145 M, Epochs=362021.07 | 45 | — | — | — | — | — | |
| ConvNeXt-SBackbone=ConvNeXt-S, Pre-training=ImageNet-1K, Params=108M, FLOPs=827G, Input Resolution=1280x800, Testing Scale=Single-scale2022.12 | 45 | 68.4 | 49.1 | — | — | — | |
| Swin-BBackbone=Swin-B, Pre-training=ImageNet-1K, Params=145M, FLOPs=982G, Input Resolution=1280x800, Testing Scale=Single-scale2022.12 | 45 | 68.4 | 48.7 | — | — | — | |
| GCNet*FLOPs=1041G, Multi-scale evaluation=false2021.07 | 44.7 | — | — | — | — | — | |
| GCNetPre-trained on=ImageNet-1K, Epochs=36, multi-scale testing=true2021.07 | 44.7 | — | — | — | — | — | |
| Swin-SBackbone=Swin-S, Pre-training=ImageNet-1K, Params=107M, FLOPs=838G, Input Resolution=1280x800, Testing Scale=Single-scale2022.12 | 44.7 | 67.9 | 48.5 | — | — | — | |
| RevCol-TBackbone=RevCol-T, Pre-training=ImageNet-1K, Params=88M, FLOPs=741G, Input Resolution=1280x800, Testing Scale=Single-scale2022.12 | 43.8 | 66.7 | 47.4 | — | — | — | |
| XCiT-M24/8Backbone=XCiT-M24/8, #params=98.9M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 43.7 | 67.5 | 46.9 | — | — | — | |
| Swin-TBackbone=Swin-T, Pre-training=ImageNet-1K, Params=86M, FLOPs=745G, Input Resolution=1280x800, Testing Scale=Single-scale2022.12 | 43.7 | 66.6 | 47.1 | — | — | — | |
| ConvNeXt-TBackbone=ConvNeXt-T, Pre-training=ImageNet-1K, Params=86M, FLOPs=741G, Input Resolution=1280x800, Testing Scale=Single-scale2022.12 | 43.7 | 66.5 | 47.3 | — | — | — | |
| Swin-SBackbone=Swin-S, #params=69.1M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 43.3 | 67.3 | 46.6 | — | — | — | |
| XCiT-S24/8Backbone=XCiT-S24/8, #params=64.5M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 43 | 66.5 | 46.1 | — | — | — | |
| XCiT-S12/8Backbone=XCiT-S12/8, #params=43.1M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 42.3 | 66 | 45.4 | — | — | — | |
| XCiT-M24/16Backbone=XCiT-M24/16, #params=101.1M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 42 | 65.6 | 44.9 | — | — | — | |
| XCiT-S24/16Backbone=XCiT-S24/16, #params=65.8M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 41.8 | 65.2 | 45 | — | — | — | |
| Swin-TBackbone=Swin-T, #params=47.8M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 41.6 | 65.1 | 44.9 | — | — | — | |
| ViL-LargeBackbone=ViL-Large, #params=76.1M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 41.3 | 64.4 | 44.5 | — | — | — | |
| XCiT-S12/16Backbone=XCiT-S12/16, #params=44.3M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 40.8 | 64 | 43.8 | — | — | — | |
| ViL-MediumBackbone=ViL-Medium, #params=60.1M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 40.7 | 63.8 | 43.7 | — | — | — | |
| PVT-LargeBackbone=PVT-Large, #params=81.0M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 40.7 | 63.4 | 43.7 | — | — | — | |
| PVT-MediumBackbone=PVT-Medium, #params=63.9M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 40.5 | 63.1 | 43.5 | — | — | — | |
| x^3 + dynamic scalePre-trained=ImageNet-1k, Model=Mask R-CNN2024.10 | 40.4 | 63.2 | 43.1 | — | — | — | |
| x^3 + fixed scalePre-trained=ImageNet-1k, Model=Mask R-CNN2024.10 | 40.2 | 63.1 | 43 | — | — | — | |
| softmaxPre-trained=ImageNet-1k, Model=Mask R-CNN2024.10 | 40.1 | 63.1 | 42.8 | — | — | — | |
| PVT-SmallBackbone=PVT-Small, #params=44.1M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 39.9 | 62.5 | 42.8 | — | — | — | |
| XCiT-T12/8Backbone=XCiT-T12/8, #params=25.8M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 39.8 | 62.2 | 43 | — | — | — | |
| ResNeXt101-64Backbone=ResNeXt101-64, #params=101.9M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 39.7 | 61.9 | 42.6 | — | — | — | |
| ViL-SmallBackbone=ViL-Small, #params=45.0M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 39.6 | 62.1 | 42.4 | — | — | — | |
| ResNeXt101-32Backbone=ResNeXt101-32, #params=62.8M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 39.2 | 61.4 | 41.9 | — | — | — | |
| ResNet101Backbone=ResNet101, #params=63.2M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 38.5 | 60.1 | 41.3 | — | — | — | |
| ViL-TinyBackbone=ViL-Tiny, #params=26.9M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 37.1 | 58.4 | 40.1 | — | — | — | |
| ResNet50Backbone=ResNet50, #params=44.2M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 37.1 | 58.4 | 40.1 | — | — | — | |
| XCiT-T12/16Backbone=XCiT-T12/16, #params=26.1M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 36.9 | 57.1 | 40 | — | — | — | |
| PVT-TinyBackbone=PVT-Tiny, #params=32.9M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 35.1 | 56.7 | 37.3 | — | — | — | |
| ResNet18Backbone=ResNet18, #params=31.2M, Detector=Mask R-CNN, Pre-training=ImageNet-1k, Training Schedule=3x2021.06 | 31.2 | 50.8 | 33.2 | — | — | — | |
| Adventurer2026.02 | — | — | — | 40.3 | 62.2 | 43.5 | |
| Copy-paste#param.=185M, FLOPs=1440G2021.03 | — | — | — | 47.2 | — | — | |
| DefMambaFLOPs (G)=2682026.02 | — | — | — | 42.8 | 66.3 | 46.2 | |
| EfficientVMamba2026.02 | — | — | — | 38.6 | 60.5 | 41.5 | |
| FractalMambaFLOPs (G)=2662026.02 | — | — | — | 42.4 | 65.9 | 45.8 | |
| GCNetMulti-scale testing=true, FLOPs=1041G2021.03 | — | — | — | 44.7 | — | — | |
| GroupMambaFLOPs (G)=2792026.02 | — | — | — | 42.9 | 66.5 | 46.3 | |
| LocalMambaFLOPs (G)=2912026.02 | — | — | — | 42.2 | 65.7 | 45.5 | |
| PlainMamba-AdapterFLOPs (G)=5422026.02 | — | — | — | 40.6 | 63.8 | 43.6 | |
| PRISMambaFLOPs (G)=2352026.02 | — | — | — | 43.2 | 67.4 | 46.8 | |
| QuadMambaFLOPs (G)=3012026.02 | — | — | — | 42.4 | 65.9 | 45.6 | |
| Swin-B (HTC++)Backbone=Swin-B, Framework=HTC++, #param.=160M, FLOPs=1043G2021.03 | — | — | — | 49.1 | — | — | |
| Swin-L (HTC++)Backbone=Swin-L, Framework=HTC++, #param.=284M, FLOPs=1470G2021.03 | — | — | — | 49.5 | — | — | |
| Swin-L (HTC++)Backbone=Swin-L, Framework=HTC++, Multi-scale testing=true, #param.=284M2021.03 | — | — | — | 50.4 | — | — | |
| Vim2026.02 | — | — | — | 39.2 | 60.9 | 41.7 | |
| VMambaFLOPs (G)=2622026.02 | — | — | — | 42.1 | 65.5 | 45.3 | |
| VSSDFLOPs (G)=2652026.02 | — | — | — | 42.6 | 66.4 | 45.9 | |
| X101-64 (HTC++)Backbone=ResNeXt-101-64x4d, Framework=HTC++, #param.=155M, FLOPs=1033G2021.03 | — | — | — | 46 | — | — |