Referring Expression Comprehension on RefCOCO (testB)
91.46AccuracyUNINEXT-H
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| UNINEXT-HModel Type=Specialist SOTAs2023.11 | 91.46 | — | — | — | — | — | — | |
| UNINEXT-HType=Specialist2024.03 | 91.46 | — | — | — | — | — | — | |
| Supervised SOTA2023.11 | 91.46 | — | — | — | — | — | — | |
| BARE-LVisual Encoder=BEiT-L [9], Fine-tuning protocol=Fine-tuning, Pre-training type=pretrained open-set detection model, Training dataset scope=mixed datasets2026.01 | 90.78 | — | — | — | — | — | — | |
| ONE-PEACEType=Specialist2024.03 | 89.26 | — | — | — | — | — | — | |
| CogVLM2023.12 | 88.73 | — | — | — | — | — | — | |
| C³VGBackbone=BEIT3-ViT-B2025.01 | 88.71 | — | — | — | — | — | — | |
| G-DINO-LType=Specialist2024.03 | 88.24 | — | — | — | — | — | — | |
| Grounding DINO-LBackbone=Swin-L, Pre-training Data=O365, OI, GoldG, Cap4M, COCO, RefC, Fine-tuning=true2023.03 | 88.24 | — | — | — | — | — | — | |
| GETok-SFTLearning Paradigm=Supervised Fine-Tuning, Evaluation Protocol=Acc@0.52025.12 | 88.2 | — | 87.2 | — | — | — | — | |
| CoT4DET-7BType=MLLM2025.12 | 88.1 | — | — | — | — | — | — | |
| GETok-SFT-gridLearning Paradigm=Supervised Fine-Tuning, Evaluation Protocol=Acc@0.52025.12 | 87.8 | — | 86.9 | — | — | — | — | |
| EEVGBackbone=ViT-B2025.01 | 87.72 | — | — | — | — | — | — | |
| SimVG-TBVisual Encoder=ViT-L/32, Time (ms)=1012024.09 | 87.68 | — | — | — | — | — | — | |
| SimVG-LVenue=NeurIPS’24, Visual Encoder=BEiT-L [9], Fine-tuning protocol=Fine-tuning, Pre-training type=pretrained vision-language model, Training dataset scope=single dataset2026.01 | 87.68 | — | — | — | — | — | — | |
| Qwen2-VL-7BType=MLLM2025.12 | 87.3 | — | — | — | — | — | — | |
| FIBER-BPre-training data (Im-Txt)=true, Pre-training data (Im-Txt-Box)=true2022.06 | 87.26 | — | — | — | — | — | — | |
| FIBERzero-shot=false2023.06 | 87.26 | — | — | — | — | — | — | |
| PolyFormer-LVisual Backbone=Swin-L, Text Encoder=BERT-base2023.02 | 87.16 | — | — | — | — | — | — | |
| SimVG-DBVisual Encoder=ViT-L/32, Time (ms)=1162024.09 | 87.07 | — | — | — | — | — | — | |
| SimVGBackbone=BEIT3-ViT-B2025.01 | 87.04 | — | — | — | — | — | — | |
| UNINEXT-LLearning Paradigm=Supervised Fine-Tuning, Evaluation Protocol=Acc@0.52025.12 | 87 | — | 88.9 | — | — | — | — | |
| CoLAUpdate Ratio=17.2%2026.04 | 86.9 | — | — | — | — | — | — | |
| ClawMachineXLearning Paradigm=Supervised Fine-Tuning, Evaluation Protocol=Acc@0.52025.12 | 86.8 | — | 86.9 | — | — | — | — | |
| MM-G-T(c3)Backbone=Swin-T, Setting=5e2024.01 | 86.6 | — | — | — | — | — | — | |
| Qwen2.5-VL-7BLearning Paradigm=Supervised Fine-Tuning, Evaluation Protocol=Acc@0.52025.12 | 86.6 | — | 85.4 | — | — | — | — | |
| GromaLearning Paradigm=Supervised Fine-Tuning, Evaluation Protocol=Acc@0.52025.12 | 86.5 | — | 86.3 | — | — | — | — | |
| Groma-7BType=MLLM2025.12 | 86.3 | — | — | — | — | — | — | |
| PolyFormer-BVisual Backbone=Swin-B, Text Encoder=BERT-base2023.02 | 86.03 | — | — | — | — | — | — | |
| PolyFormerBackbone=Swin-B2025.01 | 86.03 | — | — | — | — | — | — | |
| G-DINO-T(c)Backbone=Swin-T, Setting=fine-tuned2024.01 | 86 | — | — | — | — | — | — | |
| Grounding DINOType=Open-set Detection Model2025.12 | 86 | — | — | — | — | — | — | |
| Grounding DINO-TBackbone=Swin-T, Pre-training Data=O365, GoldG, RefC, Fine-tuning=true2023.03 | 85.99 | — | — | — | — | — | — | |
| GroundingDINOBackbone=Swin-T2025.01 | 85.99 | — | — | — | — | — | — | |
| Emu2-Chat2023.12 | 85.97 | — | — | — | — | — | — | |
| PerceptionGPT-13BModel Type=Generalist VL SOTAs2023.11 | 85.96 | — | — | — | — | — | — | |
| Qwen-VLType=Generalist, tuning=SC-Tune, Locator training stage (Lsup)=true2024.03 | 85.72 | — | — | — | — | — | — | |
| Vary-toyType=LLM-based, Size=1.8B2024.01 | 85.7 | — | — | — | — | — | — | |
| LION-12BBackbone=Flan-T5-11B2025.01 | 85.57 | — | — | — | — | — | — | |
| EEVGUpdate Ratio=100%2026.04 | 85.5 | — | — | — | — | — | — | |
| Qwen2.5-VL-7BType=MLLM2025.12 | 85.4 | — | — | — | — | — | — | |
| Qwen-VL-7B2023.12 | 85.34 | — | — | — | — | — | — | |
| OFA-LPre-training data (Im-Txt)=true, Pre-training data (Im-Txt-Box)=true2022.06 | 85.26 | — | — | — | — | — | — | |
| OFA-LVisual Backbone=RN152, Text Encoder=Embedding layer2023.02 | 85.26 | — | — | — | — | — | — | |
| OFAzero-shot=false2023.06 | 85.26 | — | — | — | — | — | — | |
| SwimVGUpdate Ratio=2.04%2026.04 | 84.9 | — | — | — | — | — | — | |
| Qwen-VLType=Generalist, tuning=SC-Tune2024.03 | 84.62 | — | — | — | — | — | — | |
| PerceptionGPT-7BModel Type=Generalist VL SOTAs2023.11 | 84.6 | — | — | — | — | — | — | |
| Qwen-VL-chatType=LLM-based, Size=7B2024.01 | 84.5 | — | — | — | — | — | — | |
| Qwen-VLType=Generalist2024.03 | 84.3 | — | — | — | — | — | — | |
| SimVG-DBVisual Encoder=ViT-B/32, Time (ms)=522024.09 | 84.04 | — | — | — | — | — | — | |
| FerretLearning Paradigm=Supervised Fine-Tuning, Evaluation Protocol=Acc@0.52025.12 | 83.9 | — | 82.5 | — | — | — | — | |
| UniTABVisual Backbone=RN101, Text Encoder=RoBERTa-base2023.02 | 83.75 | — | — | — | — | — | — | |
| SeqTRVisual Backbone=DN53, Text Encoder=Bi-GRU2023.02 | 83.59 | — | — | — | — | — | — | |
| SimVG-TBVisual Encoder=ViT-B/32, Time (ms)=442024.09 | 83.57 | — | — | — | — | — | — | |
| DQ-DETRBackbone=R101, Pre-training Data=GoldG, RefC, Fine-tuning=true2023.03 | 83.51 | — | — | — | — | — | — | |
| OFA-BVisual Backbone=RN101, Text Encoder=Embedding layer2023.02 | 83.3 | — | — | — | — | — | — | |
| HiVGUpdate Ratio=20.1%2026.04 | 83.3 | — | — | — | — | — | — | |
| VG-LAWUpdate Ratio=100%2026.04 | 83.2 | — | — | — | — | — | — | |
| UNICORN-BPre-training data (Im-Txt)=false, Pre-training data (Im-Txt-Box)=true2022.06 | 83.06 | — | — | — | — | — | — | |
| VisonLLMLearning Paradigm=Supervised Fine-Tuning, Evaluation Protocol=Acc@0.52025.12 | 82.9 | — | 80.2 | — | — | — | — | |
| ShikraLearning Paradigm=Supervised Fine-Tuning, Evaluation Protocol=Acc@0.52025.12 | 82.9 | — | 80.2 | — | — | — | — | |
| MDETRDetection backbone=ENB3, Pre-training image data=COCO, VG, Flickr30k (200k)2021.04 | 82.67 | — | — | — | — | — | — | |
| MDETR-BPre-training data (Im-Txt)=false, Pre-training data (Im-Txt-Box)=true2022.06 | 82.67 | — | — | — | — | — | — | |
| MDETRVisual Backbone=ENB3, Text Encoder=RoBERTa-base2023.02 | 82.67 | — | — | — | — | — | — | |
| MDETRzero-shot=false2023.06 | 82.67 | — | — | — | — | — | — | |
| GroundingGPTLLM Size=7B, Additional image region perception modules=false2024.01 | 82.47 | — | — | — | — | — | — | |
| FerretBackbone=Vicuna-7B2025.01 | 82.45 | — | — | — | — | — | — | |
| FerretLLM Size=7B, Additional image region perception modules=true2024.01 | 82.45 | — | — | — | — | — | — | |
| MiniGPT-v2Type=Generalist, tuning=SC-Tune, Locator training stage (Lsup)=true2024.03 | 82.33 | — | — | — | — | — | — | |
| QRNetUpdate Ratio=100%2026.04 | 82.3 | — | — | — | — | — | — | |
| Shikra-13BModel Type=Generalist VL SOTAs2023.11 | 81.81 | — | — | — | — | — | — | |
| Shikra-13B2023.12 | 81.81 | — | — | — | — | — | — | |
| Shikra-13BType=LLM-based, Size=13B2024.01 | 81.7 | — | — | — | — | — | — | |
| MDETRDetection backbone=R101, Pre-training image data=COCO, VG, Flickr30k (200k)2021.04 | 81.41 | — | — | — | — | — | — | |
| MDETRBackbone=R101, Pre-training Data=GoldG, RefC, Fine-tuning=true2023.03 | 81.41 | — | — | — | — | — | — | |
| MDETRBackbone=EfficientNet-B32025.01 | 81.41 | — | — | — | — | — | — | |
| MDETRLLM Size=-, Additional image region perception modules=false2024.01 | 81.41 | — | — | — | — | — | — | |
| SeqTRModel Type=Specialist SOTAs2023.11 | 81.24 | — | — | — | — | — | — | |
| SeqTRVisual Encoder=DN53, Time (ms)=502024.09 | 81.24 | — | — | — | — | — | — | |
| MaPPERUpdate Ratio=6.2%2026.04 | 81.2 | — | — | — | — | — | — | |
| RefTrVisual Backbone=RN101, Text Encoder=BERT-base2023.02 | 81.16 | — | — | — | — | — | — | |
| RefTRBackbone=R101, Pre-training Data=VG, Fine-tuning=true2023.03 | 81.16 | — | — | — | — | — | — | |
| MiniGPT-v2Type=Generalist2024.03 | 81.08 | — | — | — | — | — | — | |
| TransVG++Update Ratio=100%2026.04 | 81 | — | — | — | — | — | — | |
| TransVG++Backbone=ViT-B2025.01 | 80.97 | — | — | — | — | — | — | |
| InternVL2-7BType=MLLM2025.12 | 80.7 | — | — | — | — | — | — | |
| UniTABLLM Size=-, Additional image region perception modules=false2024.01 | 80.61 | — | — | — | — | — | — | |
| UniTABType=Traditional2024.01 | 80.6 | — | — | — | — | — | — | |
| Shikra-7BModel Type=Generalist VL SOTAs2023.11 | 80.24 | — | — | — | — | — | — | |
| Shikra-7BType=Generalist2024.03 | 80.24 | — | — | — | — | — | — | |
| Shikra-7B2023.12 | 80.24 | — | — | — | — | — | — | |
| Upper BoundToken=2562024.06 | 80.24 | — | — | — | — | — | — | |
| ShikraLLM Size=7B, Additional image region perception modules=false2024.01 | 80.24 | — | — | — | — | — | — | |
| MiniGPT-v2Type=Generalist, tuning=SC-Tune2024.03 | 80.22 | — | — | — | — | — | — | |
| Shikra-7BType=LLM-based, Size=7B2024.01 | 80.2 | — | — | — | — | — | — | |
| Shikra-7BType=MLLM2025.12 | 80.2 | — | — | — | — | — | — | |
| Dyn.MDETRVisual Encoder=ViT-B/162024.09 | 80.12 | — | — | — | — | — | — | |
| Dyn. MDETRBackbone=ViT-B2025.01 | 80.12 | — | — | — | — | — | — | |
| TransCPVisual Encoder=RN50, Time (ms)=74*2024.09 | 79.78 | — | — | — | — | — | — |