Referring Expression Comprehension on RefCOCOg (test)
92.2AccuracyInternVL2.5-78B
Evaluation Results
| Method | Links | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| InternVL2.5-78BParameters=78B2024.12 | 92.2 | — | — | — | — | — | — | — | — | — | — | |
| CogVLMType=Generalist2024.09 | 90.8 | — | — | — | — | — | — | — | — | — | — | |
| CogVLM-Grounding-17BParameters=17B2024.12 | 90.8 | — | — | — | — | — | — | — | — | — | — | |
| CogVLM-GroundingType=Generalist2023.11 | 90.79 | — | — | — | — | — | — | — | — | — | — | |
| CogVLM-17BModel Size=17B2024.07 | 90.79 | — | — | — | — | — | — | — | — | — | — | |
| Qwen2-VL-72BType=Generalist2024.09 | 90.4 | — | — | — | — | — | — | — | — | — | — | |
| Qwen2-VL-72BParameters=72B2024.12 | 90.4 | — | — | — | — | — | — | — | — | — | — | |
| InternVL2-26BType=Generalist2024.09 | 90.3 | — | — | — | — | — | — | — | — | — | — | |
| GETok-R1Learning Paradigm=Reinforcement Learning, Evaluation Protocol=Acc@0.52025.12 | 90.3 | — | 89.2 | — | — | — | — | — | — | — | — | |
| InternVL2-Llama3-76BBackbone=Llama3, Parameters=76B2024.12 | 90.3 | — | — | — | — | — | — | — | — | — | — | |
| VGent2025.12 | 90.1 | — | — | — | — | — | — | — | — | — | — | |
| CogVLM2023.12 | 90.09 | — | — | — | — | — | — | — | — | — | — | |
| Ferret-v2Type=Generalist2024.09 | 90 | — | — | — | — | — | — | — | — | — | — | |
| Ferret-v2-13B2025.12 | 90 | — | — | — | — | — | — | — | — | — | — | |
| InternVL3-8B2025.12 | 90 | — | — | — | — | — | — | — | — | — | — | |
| InternVL3.5-20B-A4B2025.12 | 90 | — | — | — | — | — | — | — | — | — | — | |
| Ferret-v2-13BParameters=13B2024.12 | 90 | — | — | — | — | — | — | — | — | — | — | |
| InternVL3.5-38B2025.12 | 89.9 | — | — | — | — | — | — | — | — | — | — | |
| GETok-R1-gridLearning Paradigm=Reinforcement Learning, Evaluation Protocol=Acc@0.52025.12 | 89.6 | — | 88.7 | — | — | — | — | — | — | — | — | |
| UNINEXT-HType=Specialist2024.09 | 89.4 | — | — | — | — | — | — | — | — | — | — | |
| UNINEXT-H2025.12 | 89.4 | — | — | — | — | — | — | — | — | — | — | |
| InternVL3.5-8B2025.12 | 89.4 | — | — | — | — | — | — | — | — | — | — | |
| UNINEXT-HBackbone=Huge2024.12 | 89.4 | — | — | — | — | — | — | — | — | — | — | |
| UNINEXT-HModel type=Specialist SOTAs2023.08 | 89.37 | — | — | — | — | — | — | — | — | — | — | |
| UNINEXTModel type=Specialist models2023.10 | 89.37 | — | — | — | — | — | — | — | — | — | — | |
| UNINEXT-HModel Type=Specialist SOTAs2023.11 | 89.37 | — | — | — | — | — | — | — | — | — | — | |
| UNINEXT-HType=Specialist2023.11 | 89.37 | — | — | — | — | — | — | — | — | — | — | |
| UNINEXT-HModel type=Specialist Models2023.12 | 89.37 | — | — | — | — | — | — | — | — | — | — | |
| Supervised SOTA2023.11 | 89.37 | — | — | — | — | — | — | — | — | — | — | |
| ONE-PEACEType=Specialist2024.09 | 89.3 | — | — | — | — | — | — | — | — | — | — | |
| ONE-PEACE2025.12 | 89.3 | — | — | — | — | — | — | — | — | — | — | |
| InternVL3-14B2025.12 | 89.3 | — | — | — | — | — | — | — | — | — | — | |
| ONE-PEACE2024.12 | 89.3 | — | — | — | — | — | — | — | — | — | — | |
| ONE-PEACEModel type=Specialist SOTAs2023.08 | 89.27 | — | — | — | — | — | — | — | — | — | — | |
| ONE-PEACEType=Specialist2023.11 | 89.27 | — | — | — | — | — | — | — | — | — | — | |
| CoT4DET-7BType=MLLM2025.12 | 88.9 | — | — | — | — | — | — | — | — | — | — | |
| VisionReasonerLearning Paradigm=Reinforcement Learning, Evaluation Protocol=Acc@0.52025.12 | 88.7 | — | 89 | — | — | — | — | — | — | — | — | |
| InternVL3-9B2025.12 | 88.5 | — | — | — | — | — | — | — | — | — | — | |
| LyricsModel type=Generalist Models, Backbone=Vicuna-13B2023.12 | 88.26 | — | — | — | — | — | — | — | — | — | — | |
| ASMv2-13BParameters=13B2024.02 | 88.26 | — | — | — | — | — | — | — | — | — | — | |
| GETok-SFTLearning Paradigm=Supervised Fine-Tuning, Evaluation Protocol=Acc@0.52025.12 | 88.2 | — | 88.4 | — | — | — | — | — | — | — | — | |
| Emu2-Chat2023.12 | 88.11 | — | — | — | — | — | — | — | — | — | — | |
| TextHawk22024.12 | 88.1 | — | — | — | — | — | — | — | — | — | — | |
| CoS-7BModel Size=7B, Fine-tuning protocol=LoRA2024.07 | 87.94 | — | — | — | — | — | — | — | — | — | — | |
| Qwen2-VL-7BType=Generalist2024.09 | 87.8 | — | — | — | — | — | — | — | — | — | — | |
| Qwen2-VL-7B2025.12 | 87.8 | — | — | — | — | — | — | — | — | — | — | |
| GETok-SFT-gridLearning Paradigm=Supervised Fine-Tuning, Evaluation Protocol=Acc@0.52025.12 | 87.8 | — | 87.5 | — | — | — | — | — | — | — | — | |
| Qwen2-VL-7BType=MLLM2025.12 | 87.8 | — | — | — | — | — | — | — | — | — | — | |
| Qwen2-VL-7BParameters=7B2024.12 | 87.8 | — | — | — | — | — | — | — | — | — | — | |
| InternVL2.5-8BParameters=8B2024.12 | 87.6 | — | — | — | — | — | — | — | — | — | — | |
| FIBER-BPre-training data (Im-Txt)=true, Pre-training data (Im-Txt-Box)=true2022.06 | 87.32 | — | — | — | — | — | — | — | — | — | — | |
| FIBERzero-shot=false2023.06 | 87.32 | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL-7B2025.12 | 87.2 | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL-7BType=MLLM2025.12 | 87.2 | — | — | — | — | — | — | — | — | — | — | |
| MM1.52024.12 | 87.1 | — | — | — | — | — | — | — | — | — | — | |
| G-DINO-LModel type=Specialist SOTAs2023.08 | 87.02 | — | — | — | — | — | — | — | — | — | — | |
| G-DINO-LModel type=Specialist models2023.10 | 87.02 | — | — | — | — | — | — | — | — | — | — | |
| G-DINO-LType=Specialist2023.11 | 87.02 | — | — | — | — | — | — | — | — | — | — | |
| G-DINO-LModel type=Specialist Models2023.12 | 87.02 | — | — | — | — | — | — | — | — | — | — | |
| Grounding DINO-LBackbone=Swin-L, Pre-training Data=O365, OI, GoldG, Cap4M, COCO, RefC, Fine-tuning=true2023.03 | 87.02 | — | — | — | — | — | — | — | — | — | — | |
| G-DINO-LType=Specialist2024.09 | 87 | — | — | — | — | — | — | — | — | — | — | |
| Grounding-DINO-L2025.12 | 87 | — | — | — | — | — | — | — | — | — | — | |
| UNINEXT-LLearning Paradigm=Supervised Fine-Tuning, Evaluation Protocol=Acc@0.52025.12 | 87 | — | 87.5 | — | — | — | — | — | — | — | — | |
| Groma-7BType=MLLM2025.12 | 87 | — | — | — | — | — | — | — | — | — | — | |
| Grounding-DINO-LBackbone=Large2024.12 | 87 | — | — | — | — | — | — | — | — | — | — | |
| RegionGPT2024.03 | 86.96 | — | — | — | — | — | — | — | — | — | — | |
| ClawMachineXLearning Paradigm=Supervised Fine-Tuning, Evaluation Protocol=Acc@0.52025.12 | 86.8 | — | 87.1 | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL-7BLearning Paradigm=Supervised Fine-Tuning, Evaluation Protocol=Acc@0.52025.12 | 86.6 | — | 87.2 | — | — | — | — | — | — | — | — | |
| OFA-LVisual Backbone=RN152, Text Encoder=Embedding layer2023.02 | 86.55 | — | — | — | — | — | — | — | — | — | — | |
| GromaLearning Paradigm=Supervised Fine-Tuning, Evaluation Protocol=Acc@0.52025.12 | 86.5 | — | 87 | — | — | — | — | — | — | — | — | |
| Ferret-13BEvaluation Setting=Fine-tuning Setting2023.11 | 86.34 | — | — | — | — | — | — | — | — | — | — | |
| Ferret-13BType=Generalist2023.11 | 86.34 | — | — | — | — | — | — | — | — | — | — | |
| Ferret-13BModel Size=13B2024.07 | 86.34 | — | — | — | — | — | — | — | — | — | — | |
| Ferret-13BParameters=13B2024.02 | 86.34 | — | — | — | — | — | — | — | — | — | — | |
| Qwen-VL-7B-ChatModel type=Generalist Models2023.08 | 86.32 | — | — | — | — | — | — | — | — | — | — | |
| Qwen-VLModel type=Generalist Models, Backbone=Qwen-7B2023.12 | 86.32 | — | — | — | — | — | — | — | — | — | — | |
| Qwen-VL-7BParameters=7B2024.02 | 86.32 | — | — | — | — | — | — | — | — | — | — | |
| Griffon v22024.07 | 86 | — | — | — | — | — | — | — | — | — | — | |
| PolyFormer-LVisual Backbone=Swin-L, Text Encoder=BERT-base2023.02 | 85.91 | — | — | — | — | — | — | — | — | — | — | |
| MM-G-T(c3)Backbone=Swin-T, Setting=5e2024.01 | 85.8 | — | — | — | — | — | — | — | — | — | — | |
| LION-12BEvaluation Setting=Fine-tuning Setting2023.11 | 85.74 | — | — | — | — | — | — | — | — | — | — | |
| LION-4BEvaluation Setting=Fine-tuning Setting2023.11 | 85.63 | — | — | — | — | — | — | — | — | — | — | |
| Qwen-VLType=Generalist2024.09 | 85.5 | — | — | — | — | — | — | — | — | — | — | |
| Qwen-VL-7BModel type=Generalist Models2023.08 | 85.48 | — | — | — | — | — | — | — | — | — | — | |
| Qwen-VLType=Generalist2023.11 | 85.48 | — | — | — | — | — | — | — | — | — | — | |
| Qwen-VL-7B2023.12 | 85.48 | — | — | — | — | — | — | — | — | — | — | |
| Qwen-VL-7BModel Size=7B2024.07 | 85.48 | — | — | — | — | — | — | — | — | — | — | |
| OFA-LPre-training data (Im-Txt)=true, Pre-training data (Im-Txt-Box)=true2022.06 | 85.2 | — | — | — | — | — | — | — | — | — | — | |
| OFAzero-shot=false2023.06 | 85.2 | — | — | — | — | — | — | — | — | — | — | |
| PerceptionGPT-7BModel Type=Generalist VL SOTAs2023.11 | 85.2 | — | — | — | — | — | — | — | — | — | — | |
| PolyFormer-BVisual Backbone=Swin-B, Text Encoder=BERT-base2023.02 | 84.96 | — | — | — | — | — | — | — | — | — | — | |
| Grounding DINO-TBackbone=Swin-T, Pre-training Data=O365, GoldG, RefC, Fine-tuning=true2023.03 | 84.94 | — | — | — | — | — | — | — | — | — | — | |
| G-DINO-T(c)Backbone=Swin-T, Setting=fine-tuned2024.01 | 84.9 | — | — | — | — | — | — | — | — | — | — | |
| Grounding DINOType=Open-set Detection Model2025.12 | 84.9 | — | — | — | — | — | — | — | — | — | — | |
| Ferret-7BEvaluation Setting=Fine-tuning Setting2023.11 | 84.76 | — | — | — | — | — | — | — | — | — | — | |
| Ferret (7B)Resolution=336, Visual token count=6082024.03 | 84.76 | — | — | — | — | — | — | — | — | — | — | |
| Ferret-7BModel Size=7B2024.07 | 84.76 | — | — | — | — | — | — | — | — | — | — | |
| FerretLLM Size=7B, Additional image region perception modules=true2024.01 | 84.76 | — | — | — | — | — | — | — | — | — | — | |
| UniTABVisual Backbone=RN101, Text Encoder=RoBERTa-base2023.02 | 84.7 | — | — | — | — | — | — | — | — | — | — | |
| PerceptionGPT-13BModel Type=Generalist VL SOTAs2023.11 | 84.69 | — | — | — | — | — | — | — | — | — | — |