Object Detection on COCO standard (1%)
25.81mAPLabelMatch
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| LabelMatchDetector type=Two-stage anchor-based, FLOPS=202.31G2023.02 | 25.81 | — | — | |
| Unbiased Teacher v2Thresholding Strategy=Fully automatic2022.12 | 25.4 | — | — | |
| MixPLDetector=FCOS2023.12 | 23.9 | — | — | |
| Efficient TeacherDetector type=One-stage anchor-based, Backbone=YOLOv5l, FLOPS=109.59G2023.02 | 23.76 | — | — | |
| Unbiased Teacher v2Detector type=One-stage anchor-free, FLOPS=200.59G2023.02 | 22.71 | — | — | |
| PseCoDetector type=Two-stage anchor-based, FLOPS=202.31G2023.02 | 22.43 | — | — | |
| PseCoThresholding Strategy=Fully automatic2022.12 | 22.43 | — | — | |
| Dense TeacherDetector=FCOS2023.12 | 22.4 | — | — | |
| Dense TeacherDetector type=One-stage anchor-free, FLOPS=200.59G2023.02 | 22.38 | — | — | |
| DSLDetector type=One-stage anchor-free, FLOPS=200.59G2023.02 | 22.03 | — | — | |
| DSLDetector=FCOS2023.12 | 22 | — | — | |
| Efficient TeacherDetector type=One-stage anchor-based, FLOPS=169.61G2023.02 | 21.51 | — | — | |
| Unbiased Teacher2022.03 | 20.8 | — | — | |
| Unbiased TeacherDetector type=Two-stage anchor-based, FLOPS=204.13G2023.02 | 20.75 | — | — | |
| Unbiased TeacherThresholding Strategy=Manual empirical search2022.12 | 20.75 | — | — | |
| Soft TeacherDetector type=Two-stage anchor-based, FLOPS=202.31G2023.02 | 20.46 | — | — | |
| Soft TeacherThresholding Strategy=Manual empirical search2022.12 | 20.46 | — | — | |
| ASTODThresholding Strategy=Fully automatic, Iterations=3, Refined labels=true2022.12 | 19.47 | — | — | |
| ISMTThresholding Strategy=Manual empirical search2022.12 | 18.88 | — | — | |
| Unbiased Teacher (re-implemented)Detector type=One-stage anchor-based, FLOPS=169.61G2023.02 | 18.81 | — | — | |
| Omni-DETR2022.03 | 18.6 | — | — | |
| Omni-DETRThresholding Strategy=Manual empirical search2022.12 | 18.6 | — | — | |
| Instant TeachingDetector type=Two-stage anchor-based, FLOPS=202.31G2023.02 | 18.05 | — | — | |
| Instant TeachingThresholding Strategy=Manual empirical search2022.12 | 18.05 | — | — | |
| STAC+VL-PLM2022.07 | 17.71 | — | — | |
| Humble Teacher2022.03 | 17 | — | — | |
| Humber teacherDetector type=Two-stage anchor-based, FLOPS=202.31G2023.02 | 16.96 | — | — | |
| Humble TeacherThresholding Strategy=Fully automatic2022.12 | 16.96 | — | — | |
| Supervised+VL-PLM2022.07 | 15.35 | — | — | |
| STAC2022.03 | 14 | — | — | |
| STACDetector type=Two-stage anchor-based, FLOPS=202.31G2023.02 | 13.97 | — | — | |
| STACThresholding Strategy=Manual empirical search2022.12 | 13.97 | — | — | |
| STAC2022.07 | 13.97 | — | — | |
| Supervised†Thresholding Strategy=Baseline / Teacher, Description=Our teacher model2022.12 | 12.14 | — | — | |
| Faster R-CNNImplementation=Re-implementation2022.03 | 11.7 | — | — | |
| Supervised+PLs2022.07 | 11.18 | — | — | |
| Deformable DETRImplementation=Re-implementation2022.03 | 11 | — | — | |
| CSDThresholding Strategy=Fully automatic2022.12 | 10.51 | — | — | |
| Supervised (One-stage anchor-based)Detector type=One-stage anchor-based, FLOPS=169.61G2023.02 | 10.29 | — | — | |
| Supervised (One-stage anchor-free)Detector type=One-stage anchor-free, FLOPS=200.59G2023.02 | 9.53 | — | — | |
| Supervised2022.07 | 9.25 | — | — | |
| Faster R-CNN2022.03 | 9.1 | — | — | |
| Supervised (Two-stage)Detector type=Two-stage anchor-based, FLOPS=202.31G2023.02 | 9.05 | — | — | |
| SupervisedThresholding Strategy=Baseline2022.12 | 9.05 | — | — | |
| CLIPVision Encoder=ViT-B/16, Pre-training Dataset=L-20M, Pre-training Epochs=25, Fine-tuning Epochs=122023.01 | — | — | 0.81 | |
| MAEVision Encoder=ViT-B/16, Pre-training Dataset=L-20M, Pre-training Epochs=25, Fine-tuning Epochs=122023.01 | — | — | 0.94 | |
| MAE+CLIPVision Encoder=ViT-B/16, Pre-training Dataset=L-20M, Pre-training Epochs=25, Fine-tuning Epochs=122023.01 | — | — | 0.68 | |
| MixPLDetector=DINO2023.12 | — | 31.7 | — | |
| RILSVision Encoder=ViT-B/16, Pre-training Dataset=L-20M, Pre-training Epochs=25, Fine-tuning Epochs=122023.01 | — | — | 0.86 | |
| Semi-DETRDetector=DINO2023.12 | — | 30.5 | — | |
| SLIPVision Encoder=ViT-B/16, Pre-training Dataset=L-20M, Pre-training Epochs=25, Fine-tuning Epochs=122023.01 | — | — | 1.11 |