Object Detection on COCO (test-dev)
66mAPCo-DETR
Evaluation Results
| Method | Links | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Co-DETRBackbone=ViT-L, enc. #params=304M2022.11 | 66 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DINOBackbone=InternImage-G, enc. #params=3.0B2022.11 | 65.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| M3I Pre-trainingModel=InternImage-H (1B), Pipeline=Single Stage: M3I Pre-training, Public Data=427M image-text, 15M image-category2022.11 | 65.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Focal-Stable-DINOBackbone=FocalNet-Huge [26], #Param.=689M, Pre-training Dataset=IN-22K + O365, TTA=false, w/ Mask=false2023.04 | 64.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| EVAdetector=CMask R-CNN [12], #param.=1074M, pre-training data (encoder)=merged-30M, pre-training data (detector)=O365, tta=true2022.11 | 64.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| EVAprotocol=fine-tuning2022.11 | 64.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| EVA-01Backbone=EVA, #Param.=1.0B, Pre-training Dataset=merged-30M, TTA=true, w/ Mask=true2023.04 | 64.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Group DETR v2#Params=629M, Encoder Pretraining Data=IN-1K (1M), Detector Pretraining Data=O365, w/ Mask=false2022.11 | 64.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Group DETR v2Data Constraints=Only public training data2022.11 | 64.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Group DETRv2detector=DINO [126], #param.=629M, pre-training data (encoder)=IN-1K (1M), pre-training data (detector)=O365, tta=true2022.11 | 64.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Group DETRv2protocol=fine-tuning2022.11 | 64.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Group-DETR-v2Backbone=ViT-Huge [4], #Param.=629M, Pre-training Dataset=IN-1K + O365, TTA=false, w/ Mask=false2023.04 | 64.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Co-Deformable-DETRBackbone=MixMIM-g [14], #Param.=1.0B, Pre-training Dataset=IN-1K + O365, TTA=true, w/ Mask=false2023.04 | 64.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| EVA-02Backbone=EVA-02, #Param.=304M, Pre-training Dataset=merged-38M, TTA=false, w/ Mask=true2023.04 | 64.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Group DETRv2Backbone=ViT-H, enc. #params=629M2022.11 | 64.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| EVA-02Backbone=ViT-L, enc. #params=304M2022.11 | 64.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| EVA-02-LType=Specialist Models2023.12 | 64.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| FocalNet-H (DINO)#Params.=746M, Backbone Pretraining=IN-22K (14M), Detection Pretraining=O365, W/ Mask=false, TTA=true2022.03 | 64.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| FocalNetdetector=DINO [126], #param.=746M, pre-training data (encoder)=IN-21K (14M), pre-training data (detector)=O365, tta=true2022.11 | 64.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| EVAdetector=CMask R-CNN [12], #param.=1074M, pre-training data (encoder)=merged-30M, pre-training data (detector)=O365, tta=false2022.11 | 64.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| FocalNet-DINOBackbone=FocalNet-Huge [26], #Param.=689M, Pre-training Dataset=IN-22K + O365, TTA=true, w/ Mask=false2023.04 | 64.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| EVA-01Backbone=EVA, #Param.=1.0B, Pre-training Dataset=merged-30M, TTA=false, w/ Mask=true2023.04 | 64.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| FocalNet-H (DINO)#Params=746M, Encoder Pretraining Data=IN-22K (14M), Detector Pretraining Data=O365, w/ Mask=false2022.11 | 64.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| InternImage-DINOBackbone=InternImage-XL, #Param.=602M, Pre-training Dataset=IN-22K + O365, TTA=true, w/ Mask=false2023.04 | 64.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DINOBackbone=FocalNet-H, enc. #params=746M2022.11 | 64.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| FD-SwinV2-G (HTC++)#Params=3.0B, Encoder Pretraining Data=IN-22K + IN-1K + ext-70M (85M), Detector Pretraining Data=O365, w/ Mask=true2022.11 | 64.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SwinV2-GModel=SwinV2-G (3B), Pipeline=Stage 1: Masked Image Modeling pixel, Stage 2: Image Classification, Stage 3: Dense Distillation, Public Data=15M image-category, Private Data=55M image-category2022.11 | 64.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| FD-SwinV2-Gdetector=HTC++ [17], #param.=>= 3000M, pre-training data (encoder)=IN-21K-ext-70M, pre-training data (detector)=O365, tta=null2022.11 | 64.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| FDBackbone=SwinV2-G, #Param.=3.0B, Pre-training Dataset=IN-22K-ext + O365, TTA=true, w/ Mask=true2023.04 | 64.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| FDBackbone=SwinV2-G, enc. #params=3.0B2022.11 | 64.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| FocalNet-H (DINO)#Params.=746M, Backbone Pretraining=IN-22K (14M), Detection Pretraining=O365, W/ Mask=false, TTA=false2022.03 | 64.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Stable-DINOBackbone=Swin-L, #Param.=218M, Pre-training Dataset=IN-22K + O365, TTA=false, w/ Mask=false2023.04 | 63.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| BEIT-3 (ViTDet)#Params=1.9B, Encoder Pretraining Data=IN-22K + Image-Text (35M) + Text (160GB), Detector Pretraining Data=O365, w/ Mask=true2022.11 | 63.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| BEIT-3Model=BEIT-3 (2B), Pipeline=Stage 1: CLIP, Stage 2: Dense Distillation, Stage 3: Masked Data Modeling, Public Data=21M image-text, 15M image-category, Private Data=400M image-text2022.11 | 63.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| BEIT-3detector=CMask R-CNN [12], #param.=1074M, pre-training data (encoder)=merged data^b, pre-training data (detector)=O365, tta=null2022.11 | 63.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| BEiT3Params=1.0 G, Pre-training Images=35 M, Pre-training Annotation=labeled & image-text, Detector=ViTDet, Intermediate fine-tuning (Object365)=false2022.12 | 63.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| BEiT-3Backbone=ViT-g [29], #Param.=1.9B, Pre-training Dataset=merged data^b + O365, TTA=true, w/ Mask=true2023.04 | 63.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| BEIT3Backbone=ViT-g, enc. #params=1.9B2022.11 | 63.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| RevCol-HParams=2.1 G, Pre-training Images=168 M, Pre-training Annotation=semi-labeled, Detector=DINO, Intermediate fine-tuning (Object365)=true2022.12 | 63.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DETABackbone=Swin-L, Extra Data=Objects365, FPS=4.2, Schedule=2x, Hardware=V1002022.12 | 63.5 | — | 80.4 | 70.2 | 46.1 | 66.9 | 76.9 | — | — | — | — | — | — | |
| GLIPv2detector=DyHead [28], #param.=>= 637M, pre-training data (encoder)=FLD-900M, pre-training data (detector)=merged data^a, tta=true2022.11 | 63.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DINO (Swin-L)#Params=218M, Encoder Pretraining Data=IN-22K (14M), Detector Pretraining Data=O365, w/ Mask=false2022.11 | 63.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DINO (Swin-L)#Params.=218M, Backbone Pretraining=IN-22K (14M), Detection Pretraining=O365, W/ Mask=false, TTA=true2022.03 | 63.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DINO#param.=218M, pre-training data (encoder)=IN-21K (14M), pre-training data (detector)=O365, tta=true2022.11 | 63.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DINOBackbone=Swin-L, Extra Data=Objects365, FPS=2.72022.12 | 63.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DINOBackbone=Swin-L [19], #Param.=218M, Pre-training Dataset=IN-22K + O365, TTA=false, w/ Mask=false2023.04 | 63.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DINOBackbone=Swin-L, enc. #params=218M2022.11 | 63.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DINO (Swin-L)#Params.=218M, Backbone Pretraining=IN-22K (14M), Detection Pretraining=O365, W/ Mask=false, TTA=false2022.03 | 63.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Mask Frozen-DINO-DETRbackbone=FocalNet-L, epochs=6, Object365 pre-training=true2023.08 | 63.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SwinV2-G (HTC++)train I(W) size=1536(32), test I(W) size=ms(48), multi-scale testing=true2021.11 | 63.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SwinV2-G (HTC++)#Params=3.0B, Encoder Pretraining Data=IN-22K + ext-70M (84M), Detector Pretraining Data=O365, w/ Mask=true2022.11 | 63.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SwinV2-G#Params.=3.0B, Backbone Pretraining=In-22K + ext-70M (84M), Detection Pretraining=FLD-9M, W/ Mask=true, TTA=false2022.03 | 63.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SwinV2-GModel=SwinV2-G (3B), Pipeline=Stage 1: Masked Image Modeling pixel, Stage 2: Image Classification, Public Data=15M image-category, Private Data=55M image-category2022.11 | 63.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SwinV2-Gdetector=HTC++ [17], #param.=>= 3000M, pre-training data (encoder)=IN-21K-ext-70M, pre-training data (detector)=O365, tta=true2022.11 | 63.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SwinV2-Gprotocol=fine-tuning2022.11 | 63.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SwinV2-GParams=3.0 G, Pre-training Images=70 M, Pre-training Annotation=labeled, Detector=HTC++, Intermediate fine-tuning (Object365)=false2022.12 | 63.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| HTC++Backbone=SwinV2-G [18], #Param.=3.0B, Pre-training Dataset=IN-22K-ext + O365, TTA=true, w/ Mask=true2023.04 | 63.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| HTC++Backbone=SwinV2-G, enc. #params=3.0B2022.11 | 63.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| CP-DETR-LBackbone=Swin-L, Zero-shot=true2024.12 | 62.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| FlorenceInference Mode=Fine-tuning2021.11 | 62.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Florence (CoSwin-H)#Params=≥637M, Encoder Pretraining Data=FLD-900M (900M), Detector Pretraining Data=FLD-9M, w/ Mask=true2022.11 | 62.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GLIPv2 (CoSwin-H)#Params=≥637M, Encoder Pretraining Data=FLD-900M (900M), Detector Pretraining Data=FourODs + INBoxes + GoldG + CC15M + SBU, w/ Mask=true2022.11 | 62.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Florencedetector=DyHead [28], #param.=>= 637M, pre-training data (encoder)=FLD-900M, pre-training data (detector)=merged data^a, tta=true2022.11 | 62.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Florenceprotocol=fine-tuning2022.11 | 62.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| FlorenceParams=0.9 G, Pre-training Images=900 M, Pre-training Annotation=image-text, Detector=DyHead, Intermediate fine-tuning (Object365)=false2022.12 | 62.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| FlorenceBackbone=CoSwin-H, #Param.=637M, Pre-training Dataset=FLD900M + merged data^a, TTA=false, w/ Mask=false2023.04 | 62.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GLEE-ProType=Foundation Models2023.12 | 62.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GLEE-ProBackbone=ViT-L, Zero-shot=false2024.12 | 62.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GLIP (DyHead)#Params=>284M, Encoder Pretraining Data=IN-22K (14M), Detector Pretraining Data=FourODs + GoldG + Cap24M, w/ Mask=false2022.11 | 61.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GLIPdetector=DyHead [28], #param.=>= 284M, pre-training data (encoder)=IN-21K (14M), pre-training data (detector)=4ODs+GoldG+Cap12M, tta=true2022.11 | 61.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Soft TeacherInference Mode=Fine-tuning2021.11 | 61.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SoftTeachertrain I(W) size=1280(12), test I(W) size=ms(12), multi-scale testing=true2021.11 | 61.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Soft-Teacher (Swin-L)#Params=284M, Encoder Pretraining Data=IN-22K (14M), Detector Pretraining Data=COCO-unlabeled + O365, w/ Mask=true2022.11 | 61.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Soft-Teacher (Swin-L)#Params.=284M, Backbone Pretraining=IN-22K (14M), Detection Pretraining=O365, W/ Mask=true, TTA=false2022.03 | 61.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Soft-Teacherdetector=HTC++ [17], #param.=284M, pre-training data (encoder)=IN-21K (14M), pre-training data (detector)=COCO(unlabeled)+O365, tta=true2022.11 | 61.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Soft-TeacherBackbone=Swin-L, #Param.=284M, Pre-training Dataset=IN22k + O365 + COCO(unlabeled), TTA=false, w/ Mask=true2023.04 | 61.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ViT-Adapter-LBackbone=ViT-L2022.05 | 60.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SwinV2-L (HTC++)train I(W) size=1536(32), test I(W) size=ms(48), multi-scale testing=true2021.11 | 60.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SwinV2-LBackbone=SwinV2-L2022.05 | 60.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DyHeadInference Mode=Fine-tuning2021.11 | 60.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DyHeadtrain I(W) size=1200(-), test I(W) size=ms(-), multi-scale testing=true2021.11 | 60.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DyHead (Swin-L)#Params=213M, Encoder Pretraining Data=IN-22K (14M), Detector Pretraining Data=n/a, w/ Mask=true2022.11 | 60.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GLEE-PlusType=Foundation Models2023.12 | 60.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GLIPv2-HBackbone=Swin-H, Zero-shot=false2024.12 | 60.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| CBNettrain I(W) size=1400(7), test I(W) size=ms(7), multi-scale testing=true2021.11 | 60.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| CB-Swin-L (HTC)Pre-trained on=ImageNet-22K, Params=453 M, Epochs=12, multi-scale testing=true2021.07 | 60.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| CB-Swin-LBackbone=Swin-L2022.05 | 60.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| CB-Swin-L (HTC)Pre-trained on=ImageNet-22K, Params=453 M, Epochs=122021.07 | 59.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| CB-Swin-B (HTC)Pre-trained on=ImageNet-22K, Params=235 M, Epochs=20, multi-scale testing=true2021.07 | 59.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Focal-LDetection Method=DyHead, Params=229M, Multi-scale evaluation=true2021.07 | 58.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Focal-L (DyHead)#Params.=229M, Backbone Pretraining=IN-22K (14M), Detection Pretraining=n/a, W/ Mask=true, TTA=false2022.03 | 58.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DDQ DETR 4scalesBackbone=Swin-L, Epochs=302023.03 | 58.8 | — | 77 | 64.6 | 39.4 | 62.1 | 74 | — | — | — | — | — | — | |
| Swin-LDetection Method=HTC++, Params=284M, Multi-scale evaluation=true2021.07 | 58.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Swin-LDetection Method=DyHead, Params=213M, Multi-scale evaluation=true2021.07 | 58.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SwinV1-Ltrain I(W) size=800(7), test I(W) size=ms(7), multi-scale testing=true2021.11 | 58.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Swin-L (HTC++)#Params=284M, Encoder Pretraining Data=IN-22K (14M), Detector Pretraining Data=n/a, w/ Mask=true2022.11 | 58.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Swin-L (HTC++)#Params.=284M, Backbone Pretraining=IN-22K (14M), Detection Pretraining=n/a, W/ Mask=true, TTA=true2022.03 | 58.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Swin-L (HTC++)Pre-trained on=ImageNet-22K, Params=284 M, Epochs=72, multi-scale testing=true2021.07 | 58.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| CB-Swin-B (HTC)Pre-trained on=ImageNet-22K, Params=235 M, Epochs=202021.07 | 58.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Swin-L (HTC++)Backbone=Swin-L, Framework=HTC++, Multi-scale testing=true, #param.=284M2021.03 | 58.7 | — | — | — | — | — | — | — | — | — | — | — | — |