Object Detection on LVIS
57.3APrEVA-02-L
Evaluation Results
| Method | Links | ||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| EVA-02-LType=Specialist Models2023.12 | 57.3 | — | — | 65.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| RegionCLIPVisual Encoder Pretraining Dataset=CC3M, Visual Encoder Backbone=RN50x4, Region Proposals=GT2021.12 | 50.1 | 50.1 | 51.7 | 50.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GLEE-ProType=Foundation Models2023.12 | 49.2 | — | — | 55.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GLIPv2 (H)Type=Generalist Models2023.12 | 48.9 | — | — | 60.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ViTDet-HType=Specialist Models2023.12 | 48.1 | — | — | 53.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ViTDet-LType=Specialist Models2023.12 | 46 | — | — | 51.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GLIPv2 (B)Type=Generalist Models2023.12 | 45.8 | — | — | 58.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GLEE-PlusType=Foundation Models2023.12 | 44.5 | — | — | 52.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| RegionCLIPVisual Encoder Pretraining Dataset=CC3M, Visual Encoder Backbone=RN50, Region Proposals=GT2021.12 | 40.7 | 43.5 | 47 | 44.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CLIPVisual Encoder Pretraining Dataset=CLIP400M, Visual Encoder Backbone=RN50, Region Proposals=GT2021.12 | 40.3 | 41.7 | 43.6 | 42.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GLEE-LiteType=Foundation Models2023.12 | 36.7 | — | — | 44.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| UNINEXT (H)Type=Generalist Models2023.12 | 36.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DE-ViTTotal Params=350M, Trained Params=23M, Epochs=14.42023.09 | 34.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| F-VLMTotal Params=445M, Trained Params=25M, Epochs=1182023.09 | 32.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PaLI-Xevaluation_protocol=Detection-tuned2023.05 | 31.42 | — | — | 30.64 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| OwLViT-L/16training_data=Object365 and VG datasets2023.05 | 31.2 | — | — | 34.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| OWL-VITTotal Params=433M, Trained Params=433M, Epochs=18002023.09 | 31.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SiameseIM2022.06 | 30.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| MAE2022.06 | 29.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ViLDtuning=tuned on non-rare LVIS2023.05 | 26.3 | — | — | 29.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| OwLViT-L/16tuning=tuned on non-rare LVIS2023.05 | 25.6 | — | — | 34.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| IntRec-MMbackbone=ResNet-50, modality=multimodal2026.02 | 25.6 | 34.7 | 38.2 | — | — | — | — | — | — | — | — | — | — | — | 35.4 | — | — | — | |
| MoCo-v32022.06 | 25.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CondHeadBackbone=RN50x4-C42022.12 | 25.1 | 33.4 | 37.8 | — | — | — | — | — | — | — | — | — | — | — | 33.7 | — | — | — | |
| CAKEbackbone=ResNet-502026.02 | 25 | 34.8 | 38.4 | — | — | — | — | — | — | — | — | — | — | — | 34.9 | — | — | — | |
| LBPbackbone=ResNet-502026.02 | 24.7 | 29.5 | 32.8 | — | — | — | — | — | — | — | — | — | — | — | 29.9 | — | — | — | |
| IntRec-Tbackbone=ResNet-50, modality=text-only2026.02 | 24.7 | 34.1 | 38 | — | — | — | — | — | — | — | — | — | — | — | 35 | — | — | — | |
| CoDetbackbone=ResNet-502026.02 | 24.5 | 31 | 35.4 | — | — | — | — | — | — | — | — | — | — | — | 31.7 | — | — | — | |
| OVMRbackbone=ResNet-502026.02 | 23.5 | 32.1 | 36.9 | — | — | — | — | — | — | — | — | — | — | — | 33.1 | — | — | — | |
| BARONbackbone=ResNet-502026.02 | 23.2 | 29.6 | 33.8 | — | — | — | — | — | — | — | — | — | — | — | 29.5 | — | — | — | |
| DVDetbackbone=ResNet-502026.02 | 23.1 | 31.2 | 35.4 | — | — | — | — | — | — | — | — | — | — | — | 31.2 | — | — | — | |
| RegionCLIP*Backbone=RN50x4-C4, re-evaluated with class-agnostic mask head=true2022.12 | 22.1 | 31.8 | 37 | — | — | — | — | — | — | — | — | — | — | — | 32.2 | — | — | — | |
| MICbackbone=ResNet-502026.02 | 22.1 | 33.9 | 40 | — | — | — | — | — | — | — | — | — | — | — | 33.8 | — | — | — | |
| Region-CLIPtuning=tuned on non-rare LVIS2023.05 | 22 | — | — | 32.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| RegionCLIPBackbone=RN50x4-C4, re-evaluated with class-agnostic mask head=false2022.12 | 22 | 32.1 | 36.9 | — | — | — | — | — | — | — | — | — | — | — | 32.3 | — | — | — | |
| CondHeadBackbone=RN152-FPN2022.12 | 21.9 | 28.4 | 34.6 | — | — | — | — | — | — | — | — | — | — | — | 29.7 | — | — | — | |
| VLDetbackbone=ResNet-502026.02 | 21.7 | 29.8 | 34.3 | — | — | — | — | — | — | — | — | — | — | — | 30.1 | — | — | — | |
| NRAAbackbone=ResNet-502026.02 | 21.2 | 27 | 31.5 | — | — | — | — | — | — | — | — | — | — | — | 27.8 | — | — | — | |
| DetProbackbone=ResNet-502026.02 | 20.3 | 26.5 | 28.9 | — | — | — | — | — | — | — | — | — | — | — | 26.8 | — | — | — | |
| ViLD*Backbone=RN152-FPN, re-evaluated with class-agnostic mask head=true2022.12 | 20 | 27 | 34 | — | — | — | — | — | — | — | — | — | — | — | 28.5 | — | — | — | |
| CondHeadBackbone=RN50-C42022.12 | 19.9 | 28.6 | 35.2 | — | — | — | — | — | — | — | — | — | — | — | 29.7 | — | — | — | |
| ViLDBackbone=RN152-FPN, re-evaluated with class-agnostic mask head=false2022.12 | 19.8 | 27.1 | 34.5 | — | — | — | — | — | — | — | — | — | — | — | 28.7 | — | — | — | |
| MMOVDbackbone=ResNet-502026.02 | 19.8 | 31.5 | 34.2 | — | — | — | — | — | — | — | — | — | — | — | 31.3 | — | — | — | |
| CondHeadBackbone=RN50-FPN2022.12 | 18.8 | 28.3 | 33.7 | — | — | — | — | — | — | — | — | — | — | — | 28.8 | — | — | — | |
| Deticbackbone=ResNet-502026.02 | 18.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | 27.4 | — | — | — | |
| F-VLMbackbone=ResNet-502026.02 | 18.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | 24.2 | — | — | — | |
| CCKT-Detbackbone=ResNet-502026.02 | 18.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | 27.1 | — | — | — | |
| RegionCLIPBackbone=RN50-C4, re-evaluated with class-agnostic mask head=false2022.12 | 17.1 | 27.4 | 34 | — | — | — | — | — | — | — | — | — | — | — | 28.2 | — | — | — | |
| RegionCLIPbackbone=ResNet-502026.02 | 17.1 | 27.4 | 34 | — | — | — | — | — | — | — | — | — | — | — | 28.2 | — | — | — | |
| RegionCLIP*Backbone=RN50-C4, re-evaluated with class-agnostic mask head=true2022.12 | 17 | 27.2 | 34.3 | — | — | — | — | — | — | — | — | — | — | — | 28.2 | — | — | — | |
| ViLDBackbone=RN50-FPN, re-evaluated with class-agnostic mask head=false2022.12 | 16.7 | 26.5 | 34.2 | — | — | — | — | — | — | — | — | — | — | — | 27.8 | — | — | — | |
| ViLD*Backbone=RN50-FPN, re-evaluated with class-agnostic mask head=true2022.12 | 16.6 | 27 | 33 | — | — | — | — | — | — | — | — | — | — | — | 27.5 | — | — | — | |
| ViLDbackbone=ResNet-502026.02 | 16.6 | 24.6 | 30.1 | — | — | — | — | — | — | — | — | — | — | — | 25.8 | — | — | — | |
| RegionCLIPVisual Encoder Pretraining Dataset=CC3M, Visual Encoder Backbone=RN50x4, Region Proposals=RPN2021.12 | 13.8 | 12.1 | 9.4 | 11.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PaLI-Xevaluation_protocol=Zeroshot2023.05 | 12.16 | — | — | 12.36 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CLIPVisual Encoder Pretraining Dataset=CLIP400M, Visual Encoder Backbone=RN50, Region Proposals=RPN2021.12 | 11.6 | 9.6 | 7.6 | 9.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Supervised + RFSBackbone=RN50-FPN, re-evaluated with class-agnostic mask head=false2022.12 | 11.6 | 23.5 | 32.5 | — | — | — | — | — | — | — | — | — | — | — | 24.3 | — | — | — | |
| RegionCLIPVisual Encoder Pretraining Dataset=CC3M, Visual Encoder Backbone=RN50, Region Proposals=RPN2021.12 | 10.9 | 10.4 | 8.2 | 9.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SupervisedBackbone=RN50-FPN, re-evaluated with class-agnostic mask head=false2022.12 | 3.3 | 22.5 | 34.5 | — | — | — | — | — | — | — | — | — | — | — | 23.3 | — | — | — | |
| AIMv2Evaluation protocol=Finetuning on grounding dataset mixture, Detection mode=Open-vocabulary2024.11 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 31.6 | — | — | — | |
| APE-D*training=partially trained on LVIS2025.11 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 59.6 | — | — | — | |
| BagelOutput Format=bbox2026.07 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 46.8 | |
| CLIPBackbone=ViT-B/16, Pre-training Dataset=L-20M, Pre-training Epochs=25, Fine-tuning Epochs=252023.01 | — | — | — | — | — | — | 32.3 | 48.1 | 35.1 | — | — | — | — | — | — | — | — | — | |
| CutLERBackbone=ResNet-50, Zero-shot=true, Unsupervised=true, Class-agnostic=true2023.01 | — | — | — | — | 8.4 | 21.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| CutLERSupervision=Self-supervised2024.02 | — | — | — | — | — | 23.6 | — | — | — | — | — | 13.1 | 36.2 | 55.6 | 4.5 | — | — | — | |
| DeticBackbone=ResNet-50, Training Datasets=LVIS, COCO2023.06 | — | — | — | 33 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DeticBackbone=ResNet-50, Training Datasets=LVIS, COCO, ImageNet21k, Training Data Size=12.6M2023.06 | — | — | — | 35.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DFN-CLIPEvaluation protocol=Finetuning on grounding dataset mixture, Detection mode=Open-vocabulary2024.11 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 30.7 | — | — | — | |
| DINO-X2025.11 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 52.4 | — | — | — | |
| DINOv2Evaluation protocol=Finetuning on grounding dataset mixture, Detection mode=Open-vocabulary2024.11 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 30.8 | — | — | — | |
| FreeSOLOBackbone=ResNet-101, Zero-shot=true, Unsupervised=true, Class-agnostic=true2023.01 | — | — | — | — | 3.8 | 6.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| FreeSOLOSupervision=Self-supervised2024.02 | — | — | — | — | — | 6.4 | — | — | — | — | — | 0.3 | 9.7 | 34.6 | 1.9 | — | — | — | |
| Fully-supervisedSupervision type=Fully-supervised2022.10 | — | — | — | 18.47 | — | — | — | — | — | 16.74 | 39.38 | — | — | — | — | — | — | — | |
| Gao et al.Iterations x Batch size=150K×64, Open-vocabulary=true2022.07 | — | — | — | — | 8 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| gDino-T2025.11 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 20.5 | — | 15.1 | — | |
| Gemini 2.52025.11 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 16.1 | — | |
| GiT-BBackbone=Base, Training=Universal2024.03 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 12.3 | — | — | — | |
| GiT-HBackbone=Huge, Training=Universal2024.03 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 21.7 | — | — | — | |
| GiT-LBackbone=Large, Training=Universal2024.03 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 17.3 | — | — | — | |
| GLIP (T)Training Data (SEG)=false, Training Data (DET)=true, Training Data (ITP)=false, Evaluation Protocol=Zero-shot, Backbone=Tiny2023.03 | — | — | — | 18.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GLIP-Lsetting=finetuning-free, model_size=Large2023.05 | — | — | — | 37.9 | — | — | — | — | — | 35.4 | 45.5 | — | — | — | — | — | — | — | |
| GLIP-Tsetting=finetuning-free, model_size=Tiny2023.05 | — | — | — | 26 | — | — | — | — | — | 20.8 | 42 | — | — | — | — | — | — | — | |
| Grounding DINO-Swin-TOutput Format=bbox2026.07 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 38.8 | |
| GroundingDINO-Tsetting=finetuning-free, model_size=Tiny2023.05 | — | — | — | 25.6 | — | — | — | — | — | 22.1 | 36.7 | — | — | — | — | — | — | — | |
| HASSODSupervision=Self-supervised2024.02 | — | — | — | — | — | 26.9 | — | — | — | — | — | 15.6 | 42.2 | 56.9 | 4.9 | — | — | — | |
| iFS-RCNNTested on=iFSOD & iFSIS, Supervision level=Less2022.05 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 20.76 | — | — | |
| k-meansFeature source=Cropped proposals2022.10 | — | — | — | 1.33 | — | — | — | — | — | 0.14 | 15.61 | — | — | — | — | — | — | — | |
| k-meansFeature source=FPN-based features2022.10 | — | — | — | 1.55 | — | — | — | — | — | 0.2 | 17.77 | — | — | — | — | — | — | — | |
| LLMDet-L2025.11 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 42 | — | 39.3 | — | |
| LocateAnythingOutput Format=bbox2026.07 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 50.7 | |
| MAEBackbone=ViT-B/16, Pre-training Dataset=L-20M, Pre-training Epochs=25, Fine-tuning Epochs=252023.01 | — | — | — | — | — | — | 31 | 46.2 | 33.7 | — | — | — | — | — | — | — | — | — | |
| MAE+CLIPBackbone=ViT-B/16, Pre-training Dataset=L-20M, Pre-training Epochs=25, Fine-tuning Epochs=252023.01 | — | — | — | — | — | — | 32.6 | 48.8 | 35.2 | — | — | — | — | — | — | — | — | — | |
| Mask-RCNNTested on=gFSOD & gFSIS, Supervision level=More2022.05 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 16.5 | — | — | |
| Mask+SigmoidTested on=iFSOD & iFSIS, Supervision level=Less2022.05 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 16.93 | — | — | |
| MDETRTraining Data (SEG)=false, Training Data (DET)=true, Training Data (ITP)=false, Evaluation Protocol=Zero-shot2023.03 | — | — | — | 24.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| MQ-GLIP-Lsetting=finetuning-free, model_size=Large2023.05 | — | — | — | 44 | — | — | — | — | — | 41.7 | 51.3 | — | — | — | — | — | — | — | |
| MQ-GLIP-Tsetting=finetuning-free, model_size=Tiny2023.05 | — | — | — | 30.4 | — | — | — | — | — | 26.5 | 42.8 | — | — | — | — | — | — | — | |
| MQ-GroundingDINO-Tsetting=finetuning-free, model_size=Tiny2023.05 | — | — | — | 30.2 | — | — | — | — | — | 26.2 | 43 | — | — | — | — | — | — | — | |
| OAI CLIPEvaluation protocol=Finetuning on grounding dataset mixture, Detection mode=Open-vocabulary2024.11 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | 31 | — | — | — | |
| OpenSeeD (L)Training Data (SEG)=true, Training Data (DET)=true, Training Data (ITP)=false, Evaluation Protocol=Zero-shot, Backbone=Large2023.03 | — | — | — | 23 | — | — | — | — | — | — | — | — | — | — | — | — | — | — |