Open-vocabulary Object Detection on LVIS v1 (val)
47.2AP_r^bOWL-ST+FT
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| OWL-ST+FTBackbone=SigLIP G/14, Self-training data=WebLI, Self-training vocabulary=N-grams, Human box annotations=O+VG, LVISbase, Vocabulary Access Protocol=Open vocabulary2023.06 | 47.2 | 47 | |
| OWL-ST+FTBackbone=CLIP L/14, Self-training data=WebLI, Self-training vocabulary=N-grm+curated, Human box annotations=O+VG, LVISbase, Vocabulary Access Protocol=Human-curated vocabulary2023.06 | 45.9 | 50.4 | |
| OWL-ST+FTBackbone=CLIP L/14, Self-training data=WebLI, Self-training vocabulary=N-grams, Human box annotations=O+VG, LVISbase, Vocabulary Access Protocol=Open vocabulary2023.06 | 44.6 | 49.4 | |
| OWL-ST+FTBackbone=CLIP B/16, Self-training data=WebLI, Self-training vocabulary=N-grm+curated, Human box annotations=O+VG, LVISbase, Vocabulary Access Protocol=Human-curated vocabulary2023.06 | 40.5 | 45.6 | |
| OWL-STBackbone=SigLIP G/14, Self-training data=WebLI, Self-training vocabulary=N-grams, Human box annotations=O+VG, Vocabulary Access Protocol=Open vocabulary2023.06 | 37.5 | 33.7 | |
| OWL-ST+FTBackbone=CLIP B/16, Self-training data=WebLI, Self-training vocabulary=N-grams, Human box annotations=O+VG, LVISbase, Vocabulary Access Protocol=Open vocabulary2023.06 | 36.2 | 41.8 | |
| CFM-ViTpretrained model=ViT-L/16, detector backbone=ViT-L/16, Annotation=Box AP2023.09 | 35.6 | 38.5 | |
| RO-ViTpretrained model=ViT-H/16, detector backbone=ViT-H/16, evaluation protocol=box AP2023.05 | 35.1 | 37.4 | |
| OWL-STBackbone=CLIP L/14, Self-training data=WebLI, Self-training vocabulary=N-grams, Human box annotations=O+VG, Vocabulary Access Protocol=Open vocabulary2023.06 | 34.9 | 33.5 | |
| RO-ViTpretrained model=ViT-H/16, detector backbone=ViT-H/16, evaluation protocol=mask AP2023.05 | 34.1 | 35.1 | |
| CFM-ViTpretrained model=ViT-L/16, detector backbone=ViT-L/16, Annotation=Mask AP2023.09 | 33.9 | 36.6 | |
| RO-ViTpretrained model=ViT-L/16, detector backbone=ViT-L/16, evaluation protocol=box AP2023.05 | 33.6 | 36.2 | |
| DetCLIPv2Backbone=Swin-L, Self-training data=CC15M, Self-training vocabulary=Nouns+curated, Human box annotations=O365+GoldG, Vocabulary Access Protocol=Human-curated vocabulary2023.06 | 33.3 | 36.6 | |
| F-VLMBackbone=R50x64, Human box annotations=LVISbase, Vocabulary Access Protocol=Open vocabulary2023.06 | 32.8 | 34.9 | |
| RO-ViTpretrained model=ViT-L/16, detector backbone=ViT-L/16, pre-training dataset=LAION-2B, evaluation protocol=mask AP2023.05 | 32.4 | 32.9 | |
| RO-ViTpretrained model=ViT-L/16, detector backbone=ViT-L/16, evaluation protocol=mask AP2023.05 | 32.1 | 34 | |
| RO-ViTpretrained model=ViT-L/14, detector backbone=ViT-L/14, evaluation protocol=mask AP2023.05 | 31.4 | 34 | |
| OWLBackbone=CLIP L/14, Human box annotations=O365+VG, Vocabulary Access Protocol=Open vocabulary2023.06 | 31.2 | 34.6 | |
| DetCLIPv2Backbone=Swin-T, Self-training data=CC15M, Self-training vocabulary=Nouns+curated, Human box annotations=O365+GoldG, Vocabulary Access Protocol=Human-curated vocabulary2023.06 | 31 | 32.8 | |
| 3WaysBackbone=NFNet-F6, Self-training data=CC12M, Self-training vocabulary=captions, Human box annotations=LVISbase, Vocabulary Access Protocol=Open vocabulary2023.06 | 30.1 | 44.6 | |
| CFM-ViTpretrained model=ViT-B/16, detector backbone=ViT-B/16, Annotation=Box AP2023.09 | 29.6 | 33.8 | |
| OWL-STBackbone=CLIP B/16, Self-training data=WebLI, Self-training vocabulary=N-grams, Human box annotations=O+VG, Vocabulary Access Protocol=Open vocabulary2023.06 | 29.6 | 27 | |
| CFM-ViTpretrained model=ViT-B/16, detector backbone=ViT-B/16, Annotation=Mask AP2023.09 | 28.8 | 32 | |
| RO-ViTpretrained model=ViT-B/16, detector backbone=ViT-B/16, evaluation protocol=box AP2023.05 | 28.4 | 31.9 | |
| RO-ViTpretrained model=ViT-B/16, detector backbone=ViT-B/16, evaluation protocol=mask AP2023.05 | 28 | 30.2 | |
| ViLD-Enspretrained model=EffNet-B7, detector backbone=EffNet-B7, evaluation protocol=mask AP2023.05 | 26.3 | 29.3 | |
| ViLD-Enspretrained model=EffNet-B7, detector backbone=EffNet-B7, Annotation=Mask AP2023.09 | 26.3 | 29.3 | |
| F-VLMBackbone=R50x4, Human box annotations=LVISbase, Vocabulary Access Protocol=Open vocabulary2023.06 | 26.3 | 28.5 | |
| OWL-ViTpretrained model=ViT-L/14, detector backbone=ViT-L/14, evaluation protocol=box AP2023.05 | 25.6 | 34.7 | |
| OWL-ViTpretrained model=ViT-L/14, detector backbone=ViT-L/14, Annotation=Box AP2023.09 | 25.6 | 34.7 | |
| 3WaysBackbone=NFNet-F0, Self-training data=CC12M, Self-training vocabulary=captions, Human box annotations=LVISbase, Vocabulary Access Protocol=Open vocabulary2023.06 | 25.6 | 35.7 | |
| Detic-CN2pretrained model=ViT-B/32, detector backbone=R-50, evaluation protocol=mask AP2023.05 | 24.6 | 32.4 | |
| Detic-CN2pretrained model=ViT-B/32, detector backbone=R-50, Annotation=Mask AP2023.09 | 24.6 | 32.4 | |
| DeticBackbone=R50, Self-training data=IN-21k, Self-training vocabulary=LVIS classes, Human box annotations=LVISbase, Vocabulary Access Protocol=Human-curated vocabulary2023.06 | 24.6 | 32.4 | |
| OWL-ViTpretrained model=ViT-H/14, detector backbone=ViT-H/14, evaluation protocol=box AP2023.05 | 23.3 | 35.3 | |
| OWL-ViTpretrained model=ViT-H/14, detector backbone=ViT-H/14, Annotation=Box AP2023.09 | 23.3 | 35.3 | |
| RegionCLIPpretrained model=R-50x4, detector backbone=R-50x4, evaluation protocol=mask AP2023.05 | 22 | 32.3 | |
| RegionCLIPpretrained model=R-50x4, detector backbone=R-50x4, Annotation=Mask AP2023.09 | 22 | 32.3 | |
| RegionCLIPBackbone=R50x4, Self-training data=CC3M, Self-training vocabulary=6k concepts, Human box annotations=LVISbase, Vocabulary Access Protocol=Open vocabulary2023.06 | 22 | 32.3 | |
| ViLD-Enspretrained model=ViT-L/14, detector backbone=EffNet-B7, evaluation protocol=mask AP2023.05 | 21.7 | 29.6 | |
| ViLD-Enspretrained model=ViT-L/14, detector backbone=EffNet-B7, Annotation=Mask AP2023.09 | 21.7 | 29.6 | |
| PromptDetpretrained model=ViT-B/32, detector backbone=R-50, evaluation protocol=mask AP2023.05 | 21.4 | 25.3 | |
| PromptDetpretrained model=ViT-B/32, detector backbone=R-50, Annotation=Mask AP2023.09 | 21.4 | 25.3 | |
| Rasheed et al.pretrained model=ViT-B/32, detector backbone=R-50, evaluation protocol=mask AP2023.05 | 21.1 | 25.9 | |
| Rasheed et al.pretrained model=ViT-B/32, detector backbone=R-50, Annotation=Mask AP2023.09 | 21.1 | 25.9 | |
| OWLBackbone=CLIP B/16, Human box annotations=O365+VG, Vocabulary Access Protocol=Open vocabulary2023.06 | 20.6 | 27.2 | |
| DetPro-Cascadepretrained model=ViT-B/32, detector backbone=R-50, evaluation protocol=mask AP2023.05 | 20 | 27 | |
| DetPro-Cascadepretrained model=ViT-B/32, detector backbone=R-50, Annotation=Mask AP2023.09 | 20 | 27 | |
| ViLD-Enspretrained model=ViT-B/32, detector backbone=R-152, evaluation protocol=mask AP2023.05 | 18.7 | 26 | |
| ViLD-Enspretrained model=ViT-B/32, detector backbone=R-152, Annotation=Mask AP2023.09 | 18.7 | 26 | |
| OV-DETRpretrained model=ViT-B/32, detector backbone=R-50, evaluation protocol=mask AP2023.05 | 17.4 | 26.6 | |
| OV-DETRpretrained model=ViT-B/32, detector backbone=R-50, Annotation=Mask AP2023.09 | 17.4 | 26.6 | |
| VL-PLMpretrained model=ViT-B/32, detector backbone=R-50, evaluation protocol=mask AP2023.05 | 17.2 | 27 | |
| VL-PLMpretrained model=ViT-B/32, detector backbone=R-50, Annotation=Mask AP2023.09 | 17.2 | 27 |