Referring Expression Comprehension on RefCOCO+ (testA)
91.81AccuracyCogVLM
Evaluation Results
| Method | Links | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| CogVLM2023.12 | 91.81 | — | — | — | — | — | — | — | — | — | |
| CoT4DET-7BType=MLLM2025.12 | 91.8 | — | — | — | — | — | — | — | — | — | |
| SimVG-DBVisual Encoder=ViT-L/32, Pre-train images=28K2024.09 | 91.64 | — | — | — | — | — | — | — | — | — | |
| Emu2-Chat2023.12 | 91.43 | — | — | — | — | — | — | — | — | — | |
| SimVG-TBVisual Encoder=ViT-L/32, Pre-train images=28K2024.09 | 91.02 | — | — | — | — | — | — | — | — | — | |
| C³VG2025.01 | 90.69 | — | — | — | — | — | — | — | — | — | |
| Qwen2-VL-7BType=MLLM2025.12 | 90.5 | — | — | — | — | — | — | — | — | — | |
| mPLUG-2Visual Encoder=ViT-L/142024.09 | 90.17 | — | — | — | — | — | — | — | — | — | |
| FIBER-BPre-training data (Im-Txt)=true, Pre-training data (Im-Txt-Box)=true2022.06 | 90.13 | — | — | — | — | — | — | — | — | — | |
| FIBERzero-shot=false2023.06 | 90.13 | — | — | — | — | — | — | — | — | — | |
| OFA-LPre-training data (Im-Txt)=true, Pre-training data (Im-Txt-Box)=true2022.06 | 90.1 | — | — | — | — | — | — | — | — | — | |
| OFAzero-shot=false2023.06 | 90.1 | — | — | — | — | — | — | — | — | — | |
| OFA-LVisual Encoder=RN1522024.09 | 89.87 | — | — | — | — | — | — | — | — | — | |
| PolyFormerVisual Encoder=Swin-L2024.09 | 89.77 | — | — | — | — | — | — | — | — | — | |
| UNINEXT-HModel Type=Specialist SOTAs2023.11 | 89.63 | — | — | — | — | — | — | — | — | — | |
| Supervised SOTA2023.11 | 89.63 | — | — | — | — | — | — | — | — | — | |
| SimVG-TBVisual Encoder=ViT-L/32, Time (ms)=1012024.09 | 89.61 | — | — | — | — | — | — | — | — | — | |
| LION-12B2025.01 | 89.22 | — | — | — | — | — | — | — | — | — | |
| PerceptionGPT-13BModel Type=Generalist VL SOTAs2023.11 | 89.19 | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL-7BType=MLLM2025.12 | 89.1 | — | — | — | — | — | — | — | — | — | |
| Grounding DINO-LBackbone=Swin-L, Pre-training Data=O365, OI, GoldG, Cap4M, COCO, RefC, Fine-tuning=true2023.03 | 88.95 | — | — | — | — | — | — | — | — | — | |
| Groma-7BType=MLLM2025.12 | 88.9 | — | — | — | — | — | — | — | — | — | |
| SimVG-DBVisual Encoder=ViT-B/32, Pre-train images=174K2024.09 | 88.85 | — | — | — | — | — | — | — | — | — | |
| PerceptionGPT-7BModel Type=Generalist VL SOTAs2023.11 | 88.6 | — | — | — | — | — | — | — | — | — | |
| PolyFormerVisual Encoder=Swin-B2024.09 | 88.6 | — | — | — | — | — | — | — | — | — | |
| PolyFormer2025.01 | 88.6 | — | — | — | — | — | — | — | — | — | |
| SimVG-DBVisual Encoder=ViT-B/32, Pre-train images=28K2024.09 | 88.58 | — | — | — | — | — | — | — | — | — | |
| SimVG-DBVisual Encoder=ViT-L/32, Time (ms)=1162024.09 | 88.5 | — | — | — | — | — | — | — | — | — | |
| Qwen-VL-7B2023.12 | 88.25 | — | — | — | — | — | — | — | — | — | |
| SimVG-TBVisual Encoder=ViT-B/32, Pre-train images=174K2024.09 | 88.05 | — | — | — | — | — | — | — | — | — | |
| SimVG2025.01 | 88.05 | — | — | — | — | — | — | — | — | — | |
| InternVL2-7BType=MLLM2025.12 | 87.9 | — | — | — | — | — | — | — | — | — | |
| EEVG2025.01 | 87.8 | — | — | — | — | — | — | — | — | — | |
| Shikra-13BModel Type=Generalist VL SOTAs2023.11 | 87.79 | — | — | — | — | — | — | — | — | — | |
| Shikra-13B2023.12 | 87.79 | — | — | — | — | — | — | — | — | — | |
| MM-G-T(c3)Backbone=Swin-T, Setting=5e2024.01 | 87.5 | — | — | — | — | — | — | — | — | — | |
| G-DINO-T(c)Backbone=Swin-T, Setting=fine-tuned2024.01 | 87.4 | — | — | — | — | — | — | — | — | — | |
| Grounding DINO-TBackbone=Swin-T, Pre-training Data=O365, GoldG, RefC, Fine-tuning=true2023.03 | 87.4 | — | — | — | — | — | — | — | — | — | |
| GroundingDINOVisual Encoder=Swin-T2024.09 | 87.4 | — | — | — | — | — | — | — | — | — | |
| GroundingDINO2025.01 | 87.4 | — | — | — | — | — | — | — | — | — | |
| Grounding DINOType=Open-set Detection Model2025.12 | 87.4 | — | — | — | — | — | — | — | — | — | |
| Shikra-7BType=MLLM2025.12 | 87.4 | — | — | — | — | — | — | — | — | — | |
| Ferret2025.01 | 87.38 | — | — | — | — | — | — | — | — | — | |
| FerretLLM Size=7B, Additional image region perception modules=true2024.01 | 87.38 | — | — | — | — | — | — | — | — | — | |
| Shikra-7B2023.12 | 87.36 | — | — | — | — | — | — | — | — | — | |
| Upper BoundToken=2562024.06 | 87.36 | — | — | — | — | — | — | — | — | — | |
| ShikraLLM Size=7B, Additional image region perception modules=false2024.01 | 87.36 | — | — | — | — | — | — | — | — | — | |
| Shikra-7BModel Type=Generalist VL SOTAs2023.11 | 87.25 | — | — | — | — | — | — | — | — | — | |
| GroundingGPTLLM Size=7B, Additional image region perception modules=false2024.01 | 87.18 | — | — | — | — | — | — | — | — | — | |
| VILLA_BASEdetected proposals=false, Model Scale=BASE2020.06 | 86.95 | — | — | — | — | — | — | — | — | — | |
| Aligned VLPAlignment=Supervised2022.03 | 86.6 | — | — | — | — | — | — | — | — | — | |
| UNITER_LARGEdetected proposals=false, Model Scale=LARGE2020.06 | 86.34 | — | — | — | — | — | — | — | — | — | |
| VILLA_LARGEdetected proposals=false, Model Scale=LARGE2020.06 | 86.22 | — | — | — | — | — | — | — | — | — | |
| UNITER_BASEdetected proposals=false, Model Scale=BASE2020.06 | 86.19 | — | — | — | — | — | — | — | — | — | |
| DQ-DETRBackbone=R101, Pre-training Data=GoldG, RefC, Fine-tuning=true2023.03 | 86.15 | — | — | — | — | — | — | — | — | — | |
| DQ-DETRVisual Encoder=RN1012024.09 | 86.15 | — | — | — | — | — | — | — | — | — | |
| MDETRDetection backbone=ENB3, Pre-training image data=COCO, VG, Flickr30k (200k)2021.04 | 85.52 | — | — | — | — | — | — | — | — | — | |
| MDETR-BPre-training data (Im-Txt)=false, Pre-training data (Im-Txt-Box)=true2022.06 | 85.52 | — | — | — | — | — | — | — | — | — | |
| MDETRzero-shot=false2023.06 | 85.52 | — | — | — | — | — | — | — | — | — | |
| μ-VLACCPre-training Data=CC, Alignment=Unsupervised2022.03 | 85.5 | — | — | — | — | — | — | — | — | — | |
| UniTABVisual Encoder=RN1012024.09 | 85.36 | — | — | — | — | — | — | — | — | — | |
| VoCo-LLaMAToken=82024.06 | 85.13 | — | — | — | — | — | — | — | — | — | |
| UNICORN-BPre-training data (Im-Txt)=false, Pre-training data (Im-Txt-Box)=true2022.06 | 85.05 | — | — | — | — | — | — | — | — | — | |
| μ-VLABCPre-training Data=BC, Alignment=Unsupervised2022.03 | 85 | — | — | — | — | — | — | — | — | — | |
| CoLAUpdate Ratio=17.2%2026.04 | 84.7 | — | — | — | — | — | — | — | — | — | |
| SeqTRVisual Encoder=DN532024.09 | 84.51 | — | — | — | — | — | — | — | — | — | |
| NExT-ChatLLM Size=7B, Additional image region perception modules=true2024.01 | 84.5 | — | — | — | — | — | — | — | — | — | |
| MDETRDetection backbone=R101, Pre-training image data=COCO, VG, Flickr30k (200k)2021.04 | 84.09 | — | — | — | — | — | — | — | — | — | |
| MDETRBackbone=R101, Pre-training Data=GoldG, RefC, Fine-tuning=true2023.03 | 84.09 | — | — | — | — | — | — | — | — | — | |
| MDETRVisual Encoder=RN1012024.09 | 84.09 | — | — | — | — | — | — | — | — | — | |
| MDETR2025.01 | 84.09 | — | — | — | — | — | — | — | — | — | |
| MDETRLLM Size=-, Additional image region perception modules=false2024.01 | 84.09 | — | — | — | — | — | — | — | — | — | |
| HiVGUpdate Ratio=20.1%2026.04 | 83.8 | — | — | — | — | — | — | — | — | — | |
| SimVG-TBVisual Encoder=ViT-B/32, Time (ms)=442024.09 | 83.64 | — | — | — | — | — | — | — | — | — | |
| VL-BERT_LARGEdetected proposals=false, Model Scale=LARGE2020.06 | 83.62 | — | — | — | — | — | — | — | — | — | |
| U-VisualBERTAlignment=Unsupervised2022.03 | 83.6 | — | — | — | — | — | — | — | — | — | |
| SimVG-DBVisual Encoder=ViT-B/32, Time (ms)=522024.09 | 83.36 | — | — | — | — | — | — | — | — | — | |
| UniTABLLM Size=-, Additional image region perception modules=false2024.01 | 83.22 | — | — | — | — | — | — | — | — | — | |
| SwimVGUpdate Ratio=2.04%2026.04 | 83.2 | — | — | — | — | — | — | — | — | — | |
| VoCo-LLaMAToken=12024.06 | 83.02 | — | — | — | — | — | — | — | — | — | |
| VL-BERT_BASEdetected proposals=false, Model Scale=BASE2020.06 | 82.4 | — | — | — | — | — | — | — | — | — | |
| EEVGUpdate Ratio=100%2026.04 | 82.4 | — | — | — | — | — | — | — | — | — | |
| ERNIE-ViL_LVisual Features=RN101, Pretrain Images=4.3M, Multi-task=false2021.06 | 82.37 | — | — | — | — | — | — | — | — | — | |
| Referring TransformerVisual Features=RN101, Pretrain Images=100k, Multi-task=true2021.06 | 82.26 | — | — | — | — | — | — | — | — | — | |
| RefTRBackbone=R101, Pre-training Data=VG, Fine-tuning=true2023.03 | 82.26 | — | — | — | — | — | — | — | — | — | |
| RefTRVisual Encoder=RN1012024.09 | 82.26 | — | — | — | — | — | — | — | — | — | |
| ERNIE-ViL_LDetection backbone=R101, Pre-training image data=CC, SBU (4.3M)2021.04 | 82.07 | — | — | — | — | — | — | — | — | — | |
| Dyn.MDETRVisual Encoder=ViT-B/162024.09 | 81.7 | — | — | — | — | — | — | — | — | — | |
| Dyn. MDETR2025.01 | 81.7 | — | — | — | — | — | — | — | — | — | |
| VILLA_BASEdetected proposals=true, Model Scale=BASE2020.06 | 81.65 | — | — | — | — | — | — | — | — | — | |
| VILLA_LARGEdetected proposals=true, Model Scale=LARGE2020.06 | 81.54 | — | — | — | — | — | — | — | — | — | |
| VILLA_LVisual Features=RN101, Pretrain Images=4.6M, Multi-task=false2021.06 | 81.54 | — | — | — | — | — | — | — | — | — | |
| VILLA_LDetection backbone=R101, Pre-training image data=CC, SBU, COCO, VG (4.6M)2021.04 | 81.54 | — | — | — | — | — | — | — | — | — | |
| VILLA-LPre-training data (Im-Txt)=true, Pre-training data (Im-Txt-Box)=false2022.06 | 81.54 | — | — | — | — | — | — | — | — | — | |
| VILLA_LBackbone=R101, Pre-training Data=CC, SBU, COCO, VG, Fine-tuning=true2023.03 | 81.54 | — | — | — | — | — | — | — | — | — | |
| VILLA_LVisual Encoder=RN1012024.09 | 81.54 | — | — | — | — | — | — | — | — | — | |
| UNITER_LARGEdetected proposals=true, Model Scale=LARGE2020.06 | 81.45 | — | — | — | — | — | — | — | — | — | |
| UNTIER_LVisual Features=RN101, Pretrain Images=4.6M, Multi-task=false2021.06 | 81.45 | — | — | — | — | — | — | — | — | — | |
| UNITER_LDetection backbone=R101, Pre-training image data=CC, SBU, COCO, VG (4.6M)2021.04 | 81.45 | — | — | — | — | — | — | — | — | — | |
| UNITER-LPre-training data (Im-Txt)=true, Pre-training data (Im-Txt-Box)=false2022.06 | 81.45 | — | — | — | — | — | — | — | — | — |