Object Detection on COCO novel and base categories 2014
37.4Novel AP50SAS-Det
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| SAS-DetTraining Setup=1x+Default, Backbone=RN50-C4, Detector=Faster R-CNN2023.08 | 37.4 | — | — | — | 58.5 | 53 | |
| CORATraining Setup=3x+Default, Backbone=RN50, Detector=DAB-DETR2023.08 | 35.1 | — | — | — | 35.5 | 35.4 | |
| BARONTraining Setup=1x+Default, Backbone=RN50-C4, Detector=Faster R-CNN2023.08 | 33.1 | — | — | — | 54.8 | 49.1 | |
| VL-PLMTraining Setup=1x+Default, Backbone=RN50-FPN, Detector=Faster R-CNN2023.08 | 32.3 | — | — | — | 54 | 48.3 | |
| VLDetTraining Setup=1x+LSJ, Backbone=RN50-C4, Detector=Faster R-CNN2023.08 | 32 | — | — | — | 50.6 | 45.8 | |
| PB-OVDTraining Setup=6x+Default, Backbone=RN50-C4, Detector=Faster R-CNN2023.08 | 30.8 | — | — | — | 46.1 | 42.1 | |
| OV-DETRTraining Setup=(Not Given), Backbone=RN50, Detector=Deform. DETR2023.08 | 29.4 | — | — | — | 61 | 52.7 | |
| F-VLMTraining Setup=0.5x+LSJ, Backbone=RN50-FPN, Detector=Faster R-CNN2023.08 | 28 | — | — | — | — | 39.6 | |
| DeticTraining Setup=1x+Default, Backbone=RN50-C4, Detector=Faster R-CNN2023.08 | 27.8 | — | — | — | 51.1 | 45 | |
| ViLDTraining Setup=16x+LSJ, Backbone=RN50-FPN, Detector=Faster R-CNN2023.08 | 27.6 | — | — | — | 59.5 | 51.2 | |
| RegionCLIPTraining Setup=1x+Default, Backbone=RN50-C4, Detector=Faster R-CNN2023.08 | 26.8 | — | — | — | 54.8 | 47.5 | |
| OVR-CNNTraining Setup=(Not Given), Backbone=RN50-C4, Detector=Faster R-CNN2023.08 | 22.8 | — | — | — | 46 | 39.9 | |
| Bansal et al. (2018)Training source=instance-level labels in CB, Backbone=ResNet-50, IoU threshold=0.52021.04 | — | 0.31 | 29.2 | 24.9 | — | — | |
| Bilen & Vedaldi (2016)Training source=image-level labels in CB U CN, Backbone=ResNet-50, IoU threshold=0.52021.04 | — | 19.7 | 19.6 | 19.6 | — | — | |
| CLIP on cropped regionsTraining source=image-text pairs from Internet (may contain CB U CN), Backbone=ResNet-50, IoU threshold=0.52021.04 | — | 26.3 | 28.3 | 27.8 | — | — | |
| Rahman et al. (2020)Training source=instance-level labels in CB, Backbone=ResNet-50, IoU threshold=0.52021.04 | — | 4.12 | 35.9 | 27.9 | — | — | |
| ViLD (w = 0.5)Training source=image-text pairs from Internet (may contain CB U CN), instance-level labels in CB, Backbone=ResNet-50, IoU threshold=0.5, Training iterations=450002021.04 | — | 27.6 | 59.5 | 51.3 | — | — | |
| ViLD-imageTraining source=image-text pairs from Internet (may contain CB U CN), instance-level labels in CB, Backbone=ResNet-50, IoU threshold=0.5, Training iterations=450002021.04 | — | 24.1 | 34.2 | 31.6 | — | — | |
| ViLD-textTraining source=image-text pairs from Internet (may contain CB U CN), instance-level labels in CB, Backbone=ResNet-50, IoU threshold=0.5, Training iterations=450002021.04 | — | 5.9 | 61.8 | 47.2 | — | — | |
| Ye et al. (2019)Training source=image-level labels in CB U CN, Backbone=ResNet-50, IoU threshold=0.52021.04 | — | 20.3 | 20.1 | 20.1 | — | — | |
| Zareian et al. (2021)Training source=image captions in CB U CN, instance-level labels in CB, Backbone=ResNet-50, IoU threshold=0.52021.04 | — | 22.8 | 46 | 39.9 | — | — | |
| Zhu et al. (2020)Training source=instance-level labels in CB, Backbone=ResNet-50, IoU threshold=0.52021.04 | — | 3.41 | 13.8 | 13 | — | — |