Object Detection on COCO mini 2017 (val)
43.2mAPNASNet-A (6 @ 4032)
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| NASNet-A (6 @ 4032)resolution=1200 x 1200, detection_framework=Faster-RCNN, crop=single2017.07 | 43.2 | — | — | — | — | — | — | |
| Teacher w. ResNet-101 (3x)Detector=Faster R-CNN, Backbone=ResNet-101, Training Schedule=3x2021.10 | 42 | — | — | — | 25.2 | 45.6 | 54.6 | |
| TNT-SBackbone=TNT-S, Params (M)=48.1, Epochs=12, Pre-training=ImageNet, Framework=Faster R-CNN with FPN2021.02 | 41.5 | 64.1 | 44.5 | — | 25.7 | 44.6 | 55.4 | |
| NASNet-A (6 @ 4032)resolution=800 x 800, detection_framework=Faster-RCNN, crop=single2017.07 | 41.3 | — | — | — | — | — | — | |
| ICDDetector=Faster R-CNN, Backbone=ResNet-50, Training Schedule=1x, Distillation Mode=Distilled from ResNet-101, Inheriting strategy=true2021.10 | 40.9 | — | — | — | 24.5 | 44.2 | 53.5 | |
| ICDDetector=RetinaNet, Backbone=ResNet-50, Training Schedule=1x, Distillation Mode=Distilled from ResNet-101, Inheriting strategy=true2021.10 | 40.7 | — | — | — | 24.2 | 45 | 52.7 | |
| ICDDetector=Faster R-CNN, Backbone=ResNet-50, Training Schedule=1x, Distillation Mode=Distilled from ResNet-1012021.10 | 40.4 | — | — | — | 23.4 | 44 | 52 | |
| Teacher w. ResNet-101 (3x)Detector=RetinaNet, Backbone=ResNet-101, Training Schedule=3x2021.10 | 40.4 | — | — | — | 24 | 44.3 | 52.2 | |
| DRConv-MaskRCNN 16RFramework=Mask R-CNN, Backbone=ResNet-50, Neck=FPN with DRConv, Region Number=16R2020.03 | 40.3 | 61.2 | 44.2 | — | — | — | — | |
| DRConv-MaskRCNN 8RFramework=Mask R-CNN, Backbone=ResNet-50, Neck=FPN with DRConv, Region Number=8R2020.03 | 40.2 | 60.8 | 44 | — | — | — | — | |
| Zhang et al.Detector=Faster R-CNN, Backbone=ResNet-50, Training Schedule=1x, Distillation Mode=Distilled from ResNet-1012021.10 | 40 | — | — | — | 23.2 | 43.3 | 52.5 | |
| DeiT-SBackbone=DeiT-S, Params (M)=46.4, Epochs=12, Pre-training=ImageNet, Framework=Faster R-CNN with FPN, Implementation=Ours2021.02 | 39.9 | 62.8 | 42.6 | — | 23.4 | 42.5 | 54 | |
| ICDDetector=RetinaNet, Backbone=ResNet-50, Training Schedule=1x, Distillation Mode=Distilled from ResNet-1012021.10 | 39.9 | — | — | — | 25 | 43.9 | 51 | |
| DRConv-MaskRCNN 4RFramework=Mask R-CNN, Backbone=ResNet-50, Neck=FPN with DRConv, Region Number=4R2020.03 | 39.8 | 60.3 | 43.3 | — | — | — | — | |
| FRNimgs/gpu=8, Training protocol=From scratch, Backbone=ResNetV1-101 FPN, Input resolution=1024x10242019.11 | 39.6 | 58.5 | 43.1 | — | — | — | — | |
| FRNimgs/gpu=4, Training protocol=From scratch, Backbone=ResNetV1-101 FPN, Input resolution=1024x10242019.11 | 39.5 | 58.4 | 43.3 | — | — | — | — | |
| Li et al.Detector=Faster R-CNN, Backbone=ResNet-50, Training Schedule=1x, Distillation Mode=Distilled from ResNet-1012021.10 | 39.5 | — | — | — | 23.3 | 43 | 51.4 | |
| GNimgs/gpu=8, Training protocol=From scratch, Backbone=ResNetV1-101 FPN, Input resolution=1024x10242019.11 | 39.3 | 57.8 | 42.6 | — | — | — | — | |
| FitNetDetector=Faster R-CNN, Backbone=ResNet-50, Training Schedule=1x, Distillation Mode=Distilled from ResNet-1012021.10 | 39.3 | — | — | — | 22.7 | 42.3 | 51.7 | |
| Zhang et al.Detector=RetinaNet, Backbone=ResNet-50, Training Schedule=1x, Distillation Mode=Distilled from ResNet-1012021.10 | 39.3 | — | — | — | 23.4 | 43.6 | 50.6 | |
| Wang et al.Detector=Faster R-CNN, Backbone=ResNet-50, Training Schedule=1x, Distillation Mode=Distilled from ResNet-1012021.10 | 39.2 | — | — | — | 23.2 | 42.8 | 50.4 | |
| FRNimgs/gpu=2, Training protocol=From scratch, Backbone=ResNetV1-101 FPN, Input resolution=1024x10242019.11 | 39.1 | 57.5 | 42.3 | — | — | — | — | |
| MaskRCNNFramework=Mask R-CNN, Backbone=ResNet-50, Neck=FPN2020.03 | 39.1 | 59 | 42.8 | — | — | — | — | |
| GNimgs/gpu=4, Training protocol=From scratch, Backbone=ResNetV1-101 FPN, Input resolution=1024x10242019.11 | 39 | 57.5 | 42.3 | — | — | — | — | |
| BNimgs/gpu=8, Training protocol=From scratch, Backbone=ResNetV1-101 FPN, Input resolution=1024x10242019.11 | 38.7 | 56.6 | 42.1 | — | — | — | — | |
| GNimgs/gpu=2, Training protocol=From scratch, Backbone=ResNetV1-101 FPN, Input resolution=1024x10242019.11 | 38.7 | 56.9 | 41.8 | — | — | — | — | |
| DRConv-DetNAS-300M 8RBackbone=DetNAS-300M, Region Number=8R2020.03 | 38.4 | 59.6 | 41.6 | — | — | — | — | |
| Wang et al.Detector=RetinaNet, Backbone=ResNet-50, Training Schedule=1x, Distillation Mode=Distilled from ResNet-1012021.10 | 38.4 | — | — | — | 23.3 | 42.6 | 49.1 | |
| BN*imgs/gpu=8, Training protocol=Fine-tuned, Backbone=ResNetV1-101 FPN, Input resolution=1024x10242019.11 | 38.3 | 57.2 | 41.5 | — | — | — | — | |
| FitNetDetector=RetinaNet, Backbone=ResNet-50, Training Schedule=1x, Distillation Mode=Distilled from ResNet-1012021.10 | 38.2 | — | — | — | 21.8 | 42.6 | 48.8 | |
| BNimgs/gpu=4, Training protocol=From scratch, Backbone=ResNetV1-101 FPN, Input resolution=1024x10242019.11 | 37.9 | 55.2 | 41.4 | — | — | — | — | |
| Student w. ResNet-50 (1x)Detector=Faster R-CNN, Backbone=ResNet-50, Training Schedule=1x2021.10 | 37.9 | — | — | — | 22.4 | 41.1 | 49.1 | |
| ResNet-50Backbone=ResNet-50, Params (M)=41.5, Epochs=12, Pre-training=ImageNet, Framework=Faster R-CNN with FPN2021.02 | 37.4 | 58.1 | 40.4 | — | 21.2 | 41 | 48.1 | |
| Student w. ResNet-50 (1x)Detector=RetinaNet, Backbone=ResNet-50, Training Schedule=1x2021.10 | 37.4 | — | — | — | 23.1 | 41.6 | 48.3 | |
| Inception-ResNet-v2 (TDM)resolution=600 x 1000, detection_framework=Faster-RCNN, crop=single2017.07 | 37.3 | — | — | — | — | — | — | |
| BN*imgs/gpu=4, Training protocol=Fine-tuned, Backbone=ResNetV1-101 FPN, Input resolution=1024x10242019.11 | 37.1 | 55.4 | 40.4 | — | — | — | — | |
| DetNAS-300MBackbone=DetNAS-300M2020.03 | 36.6 | 57.4 | 39.3 | — | — | — | — | |
| Inception-ResNet-v2 (G-RMI)resolution=600 x 600, detection_framework=Faster-RCNN, crop=single2017.07 | 35.7 | — | — | — | — | — | — | |
| BN*imgs/gpu=2, Training protocol=Fine-tuned, Backbone=ResNetV1-101 FPN, Input resolution=1024x10242019.11 | 32.9 | 49.1 | 35.9 | — | — | — | — | |
| BNimgs/gpu=2, Training protocol=From scratch, Backbone=ResNetV1-101 FPN, Input resolution=1024x10242019.11 | 30.2 | 44.5 | 32.5 | — | — | — | — | |
| NASNet-A (4 @ 1056)resolution=600 x 600, detection_framework=Faster-RCNN, crop=single2017.07 | 29.6 | — | — | — | — | — | — | |
| MobileNetV2 1.0xDetection Framework=Faster R-CNN, Backbone FLOPs=300M2019.11 | 27.5 | — | — | — | — | — | — | |
| MobileNetV3 1.0xDetection Framework=Faster R-CNN, Backbone FLOPs=219M2019.11 | 26.9 | — | — | — | — | — | — | |
| GhostNet 1.1xDetection Framework=Faster R-CNN, Backbone FLOPs=164M2019.11 | 26.9 | — | — | — | — | — | — | |
| MobileNetV2 1.0xDetection Framework=RetinaNet, Backbone FLOPs=300M2019.11 | 26.7 | — | — | — | — | — | — | |
| GhostNet 1.1xDetection Framework=RetinaNet, Backbone FLOPs=164M2019.11 | 26.6 | — | — | — | — | — | — | |
| MobileNetV3 1.0xDetection Framework=RetinaNet, Backbone FLOPs=219M2019.11 | 26.4 | — | — | — | — | — | — | |
| ShuffleNet (2x)resolution=600 x 600, detection_framework=Faster-RCNN, crop=single2017.07 | 24.5 | — | — | — | — | — | — | |
| MobileNet-224resolution=600 x 600, detection_framework=Faster-RCNN, crop=single2017.07 | 19.8 | — | — | — | — | — | — | |
| BYOLPre-training Epochs=3002021.02 | — | — | — | 20.6 | — | — | — | |
| InfoMinPre-training Epochs=2002021.02 | — | — | — | 23.6 | — | — | — | |
| InsLocPre-training Epochs=2002021.02 | — | — | — | 26 | — | — | — | |
| MoCo-v2Pre-training Epochs=2002021.02 | — | — | — | 22.7 | — | — | — | |
| Relative Loc.Pre-training Epochs=2002021.02 | — | — | — | 17.2 | — | — | — | |
| SimCLRPre-training Epochs=2002021.02 | — | — | — | 20 | — | — | — | |
| SupervisedPre-training Epochs=902021.02 | — | — | — | 22.9 | — | — | — | |
| SwAVPre-training Epochs=4002021.02 | — | — | — | 14.9 | — | — | — |