Referring Expression Comprehension on RefCOCO+ (val)
90.4AccuracyInternVL2.5-78B
Evaluation Results
| Method | Links | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| InternVL2.5-78BParameters=78B2024.12 | 90.4 | — | — | — | — | — | — | — | — | — | — | |
| Qwen2-VL-72BType=Generalist2024.09 | 90.1 | — | — | — | — | — | — | — | — | — | — | |
| Qwen2-VL-72BParameters=72B2024.12 | 90.1 | — | — | — | — | — | — | — | — | — | — | |
| ChatRexParameters=7B2026.02 | 89.8 | — | — | — | — | — | — | — | — | — | — | |
| InternVL2-26BType=Generalist2024.09 | 88.8 | — | — | — | — | — | — | — | — | — | — | |
| ONE-PEACEType=Specialist2024.09 | 88.8 | — | — | — | — | — | — | — | — | — | — | |
| ONE-PEACE2025.12 | 88.8 | — | — | — | — | — | — | — | — | — | — | |
| ONE-PEACE2024.12 | 88.8 | — | — | — | — | — | — | — | — | — | — | |
| InternVL2-Llama3-76BBackbone=Llama3, Parameters=76B2024.12 | 88.8 | — | — | — | — | — | — | — | — | — | — | |
| ONE-PEACEModel type=Specialist SOTAs (Specialist/Finetuned)2023.06 | 88.77 | — | — | — | — | — | — | — | — | — | — | |
| ONE-PEACEModel type=Specialist SOTAs2023.08 | 88.77 | — | — | — | — | — | — | — | — | — | — | |
| ONE-PEACEType=Specialist2023.11 | 88.77 | — | — | — | — | — | — | — | — | — | — | |
| CogVLMType=Generalist2024.09 | 88.7 | — | — | — | — | — | — | — | — | — | — | |
| CogVLM-Grounding-17BParameters=17B2024.12 | 88.7 | — | — | — | — | — | — | — | — | — | — | |
| CogVLM-GroundingType=Generalist2023.11 | 88.68 | — | — | — | — | — | — | — | — | — | — | |
| CogVLM-17BModel Size=17B2024.07 | 88.68 | — | — | — | — | — | — | — | — | — | — | |
| BARE-LVisual Encoder=BEiT-L [9], Fine-tuning protocol=Fine-tuning, Pre-training type=pretrained open-set detection model, Training dataset scope=mixed datasets2026.01 | 88.36 | — | — | — | — | — | — | — | — | — | — | |
| InternVL3-8B2025.12 | 88.2 | — | — | — | — | — | — | — | — | — | — | |
| VGent2025.12 | 88.1 | — | — | — | — | — | — | — | — | — | — | |
| InternVL3.5-8B2025.12 | 87.9 | — | — | — | — | — | — | — | — | — | — | |
| InternVL3.5-20B-A4B2025.12 | 87.6 | — | — | — | — | — | — | — | — | — | — | |
| CogVLM2023.12 | 87.52 | — | — | — | — | — | — | — | — | — | — | |
| InternVL3.5Parameters=38B2026.02 | 87.5 | — | — | — | — | — | — | — | — | — | — | |
| InternVL3.5-38B2025.12 | 87.5 | — | — | — | — | — | — | — | — | — | — | |
| C³VG2025.01 | 87.44 | — | — | — | — | — | — | — | — | — | — | |
| Ferret-v2Type=Generalist2024.09 | 87.4 | — | — | — | — | — | — | — | — | — | — | |
| Ferret-v2-13B2025.12 | 87.4 | — | — | — | — | — | — | — | — | — | — | |
| InternVL3-14B2025.12 | 87.4 | — | — | — | — | — | — | — | — | — | — | |
| Ferret-v2-13BParameters=13B2024.12 | 87.4 | — | — | — | — | — | — | — | — | — | — | |
| Emu2-Chat2023.12 | 87.05 | — | — | — | — | — | — | — | — | — | — | |
| CoT4DET-7BType=MLLM2025.12 | 86.5 | — | — | — | — | — | — | — | — | — | — | |
| VLM-FO1Parameters=3B2026.02 | 86.4 | — | — | — | — | — | — | — | — | — | — | |
| InternVL3-9B2025.12 | 86.4 | — | — | — | — | — | — | — | — | — | — | |
| TextHawk22024.12 | 86.2 | — | — | — | — | — | — | — | — | — | — | |
| ObjEmbedParameters=4B2026.02 | 86.1 | — | — | — | — | — | — | — | — | — | — | |
| CoS-7BModel Size=7B, Fine-tuning protocol=LoRA2024.07 | 86.03 | — | — | — | — | — | — | — | — | — | — | |
| OFA-LVisual Backbone=RN152, Text Encoder=Embedding layer2023.02 | 85.8 | — | — | — | — | — | — | — | — | — | — | |
| Qwen2-VL-7BType=Generalist2024.09 | 85.8 | — | — | — | — | — | — | — | — | — | — | |
| Qwen2-VL-7B2025.12 | 85.8 | — | — | — | — | — | — | — | — | — | — | |
| Qwen2-VL-7BType=MLLM2025.12 | 85.8 | — | — | — | — | — | — | — | — | — | — | |
| Qwen2-VL-7BParameters=7B2024.12 | 85.8 | — | — | — | — | — | — | — | — | — | — | |
| FIBER-BPre-training data (Im-Txt)=true, Pre-training data (Im-Txt-Box)=true2022.06 | 85.74 | — | — | — | — | — | — | — | — | — | — | |
| FIBERzero-shot=false2023.06 | 85.74 | — | — | — | — | — | — | — | — | — | — | |
| SimVG-TBVisual Encoder=ViT-L/32, Time (ms)=1012024.09 | 85.36 | — | — | — | — | — | — | — | — | — | — | |
| UNINEXT-HModel type=Specialist SOTAs (Specialist/Finetuned)2023.06 | 85.24 | — | — | — | — | — | — | — | — | — | — | |
| UNINEXT-HModel type=Specialist SOTAs2023.08 | 85.24 | — | — | — | — | — | — | — | — | — | — | |
| UNINEXTModel type=Specialist models2023.10 | 85.24 | — | — | — | — | — | — | — | — | — | — | |
| UNINEXT-HModel Type=Specialist SOTAs2023.11 | 85.24 | — | — | — | — | — | — | — | — | — | — | |
| UNINEXT-HType=Specialist2023.11 | 85.24 | — | — | — | — | — | — | — | — | — | — | |
| UNINEXT-HModel type=Specialist Models2023.12 | 85.24 | — | — | — | — | — | — | — | — | — | — | |
| Supervised SOTA2023.11 | 85.24 | — | — | — | — | — | — | — | — | — | — | |
| UNINEXT-HType=Specialist2024.09 | 85.2 | — | — | — | — | — | — | — | — | — | — | |
| InternVL2.5Parameters=8B2026.02 | 85.2 | — | — | — | — | — | — | — | — | — | — | |
| UNINEXT-H2025.12 | 85.2 | — | — | — | — | — | — | — | — | — | — | |
| UNINEXT-HBackbone=Huge2024.12 | 85.2 | — | — | — | — | — | — | — | — | — | — | |
| InternVL2.5-8BParameters=8B2024.12 | 85.2 | — | — | — | — | — | — | — | — | — | — | |
| PolyFormer-LVisual Backbone=Swin-L, Text Encoder=BERT-base2023.02 | 84.98 | — | — | — | — | — | — | — | — | — | — | |
| SimVG-DBVisual Encoder=ViT-L/32, Time (ms)=1162024.09 | 84.88 | — | — | — | — | — | — | — | — | — | — | |
| ASMv2-13BParameters=13B2024.02 | 84.81 | — | — | — | — | — | — | — | — | — | — | |
| OFA-LPre-training data (Im-Txt)=true, Pre-training data (Im-Txt-Box)=true2022.06 | 84.49 | — | — | — | — | — | — | — | — | — | — | |
| OFAzero-shot=false2023.06 | 84.49 | — | — | — | — | — | — | — | — | — | — | |
| VILLA_LARGEdetected proposals=false, Model Scale=LARGE2020.06 | 84.4 | — | — | — | — | — | — | — | — | — | — | |
| VILLA_BASEdetected proposals=false, Model Scale=BASE2020.06 | 84.26 | — | — | — | — | — | — | — | — | — | — | |
| UNITER_LARGEdetected proposals=false, Model Scale=LARGE2020.06 | 84.25 | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VLParameters=7B2026.02 | 84.2 | — | — | — | — | — | — | — | — | — | — | |
| VLM-R1Parameters=3B2026.02 | 84.2 | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL-7B2025.12 | 84.2 | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL-7BType=MLLM2025.12 | 84.2 | — | — | — | — | — | — | — | — | — | — | |
| LION-12BEvaluation Setting=Fine-tuning Setting2023.11 | 83.95 | — | — | — | — | — | — | — | — | — | — | |
| LION-12B2025.01 | 83.95 | — | — | — | — | — | — | — | — | — | — | |
| Groma-7BType=MLLM2025.12 | 83.9 | — | — | — | — | — | — | — | — | — | — | |
| PolyFormer-BVisual Backbone=Swin-B, Text Encoder=BERT-base2023.02 | 83.73 | — | — | — | — | — | — | — | — | — | — | |
| PolyFormer2025.01 | 83.73 | — | — | — | — | — | — | — | — | — | — | |
| PerceptionGPT-13BModel Type=Generalist VL SOTAs2023.11 | 83.72 | — | — | — | — | — | — | — | — | — | — | |
| UNITER_BASEdetected proposals=false, Model Scale=BASE2020.06 | 83.66 | — | — | — | — | — | — | — | — | — | — | |
| LION-4BEvaluation Setting=Fine-tuning Setting2023.11 | 83.6 | — | — | — | — | — | — | — | — | — | — | |
| OctopusParameters=7B2026.02 | 83.6 | — | — | — | — | — | — | — | — | — | — | |
| SimVG2025.01 | 83.54 | — | — | — | — | — | — | — | — | — | — | |
| ObjEmbedParameters=2B2026.02 | 83.4 | — | — | — | — | — | — | — | — | — | — | |
| BARE-BVisual Encoder=BEiT-B [9], Fine-tuning protocol=Fine-tuning, Pre-training type=pretrained open-set detection model, Training dataset scope=mixed datasets2026.01 | 83.14 | — | — | — | — | — | — | — | — | — | — | |
| Qwen-VL-7BModel type=Generalist Models2023.08 | 83.12 | — | — | — | — | — | — | — | — | — | — | |
| Qwen-VLType=Generalist2023.11 | 83.12 | — | — | — | — | — | — | — | — | — | — | |
| Qwen-VL-7B2023.12 | 83.12 | — | — | — | — | — | — | — | — | — | — | |
| Qwen-VL-7BModel Size=7B2024.07 | 83.12 | — | — | — | — | — | — | — | — | — | — | |
| Qwen-VLType=Generalist2024.09 | 83.1 | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-VLParameters=4B2026.02 | 82.9 | — | — | — | — | — | — | — | — | — | — | |
| Shikra-13BModel type=Generalist VL SOTAs (w/o finetuning)2023.06 | 82.89 | — | — | — | — | — | — | — | — | — | — | |
| Shikra-13BModel type=Generalist Models2023.08 | 82.89 | — | — | — | — | — | — | — | — | — | — | |
| Shikra (13B)Model type=Generalist models2023.10 | 82.89 | — | — | — | — | — | — | — | — | — | — | |
| Shikra-13BModel Type=Generalist VL SOTAs2023.11 | 82.89 | — | — | — | — | — | — | — | — | — | — | |
| Shikra-13BEvaluation Setting=Fine-tuning Setting2023.11 | 82.89 | — | — | — | — | — | — | — | — | — | — | |
| Shikra-13BType=Generalist2023.11 | 82.89 | — | — | — | — | — | — | — | — | — | — | |
| Shikra (13B)Resolution=224, Visual token count=2562024.03 | 82.89 | — | — | — | — | — | — | — | — | — | — | |
| ShikraModel type=Generalist Models, Backbone=Vicuna-13B2023.12 | 82.89 | — | — | — | — | — | — | — | — | — | — | |
| LyricsModel type=Generalist Models, Backbone=Vicuna-13B2023.12 | 82.89 | — | — | — | — | — | — | — | — | — | — | |
| Shikra-13B2023.12 | 82.89 | — | — | — | — | — | — | — | — | — | — | |
| Shikra-13BModel Size=13B2024.07 | 82.89 | — | — | — | — | — | — | — | — | — | — | |
| Shikra-13BParameters=13B2024.02 | 82.89 | — | — | — | — | — | — | — | — | — | — | |
| Elysium (7B)Resolution=336, Visual token count=1082024.03 | 82.86 | — | — | — | — | — | — | — | — | — | — | |
| Qwen-VL-7B-ChatModel type=Generalist Models2023.08 | 82.82 | — | — | — | — | — | — | — | — | — | — |