Referring Expression Comprehension on RefCOCO+ (test-A)
94.7AccuracyInternVL2.5-78B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| InternVL2.5-78BParameters=78B2024.12 | 94.7 | — | — | |
| Qwen2-VL-72BType=Generalist2024.09 | 93.8 | — | — | |
| Qwen2-VL-72BParameters=72B2024.12 | 93.8 | — | — | |
| InternVL2-26BType=Generalist2024.09 | 93.1 | — | — | |
| InternVL2-Llama3-76BBackbone=Llama3, Parameters=76B2024.12 | 93.1 | — | — | |
| CogVLM-GroundingType=Generalist2023.11 | 92.91 | — | — | |
| CogVLM-17BModel Size=17B2024.07 | 92.91 | — | — | |
| CogVLMType=Generalist2024.09 | 92.9 | — | — | |
| CogVLM-Grounding-17BParameters=17B2024.12 | 92.9 | — | — | |
| InternVL3-8B2025.12 | 92.5 | — | — | |
| InternVL3.5-8B2025.12 | 92.4 | — | — | |
| ONE-PEACEModel type=Specialist SOTAs (Specialist/Finetuned)2023.06 | 92.21 | — | — | |
| ONE-PEACEModel type=Specialist SOTAs2023.08 | 92.21 | — | — | |
| ONE-PEACEType=Specialist2023.11 | 92.21 | — | — | |
| ONE-PEACEType=Specialist2024.09 | 92.2 | — | — | |
| ONE-PEACE2025.12 | 92.2 | — | — | |
| VGent2025.12 | 92.2 | — | — | |
| ONE-PEACE2024.12 | 92.2 | — | — | |
| Ferret-v2Type=Generalist2024.09 | 92.1 | — | — | |
| Ferret-v2-13B2025.12 | 92.1 | — | — | |
| InternVL3-14B2025.12 | 92.1 | — | — | |
| Ferret-v2-13BParameters=13B2024.12 | 92.1 | — | — | |
| InternVL3.5-20B-A4B2025.12 | 92 | — | — | |
| InternVL2.5-8BParameters=8B2024.12 | 91.5 | — | — | |
| CoS-7BModel Size=7B, Fine-tuning protocol=LoRA2024.07 | 91.02 | — | — | |
| InternVL3-9B2025.12 | 91 | — | — | |
| ASMv2-13BParameters=13B2024.02 | 90.83 | — | — | |
| Qwen2-VL-7BType=Generalist2024.09 | 90.5 | — | — | |
| Qwen2-VL-7B2025.12 | 90.5 | — | — | |
| Qwen2-VL-7BParameters=7B2024.12 | 90.5 | — | — | |
| GETok-R1Learning Paradigm=Reinforcement Learning, Evaluation Protocol=Acc@0.52025.12 | 90.3 | 90.8 | — | |
| InternVL3.5-38B2025.12 | 90 | — | — | |
| TextHawk22024.12 | 90 | — | — | |
| OFA-LVisual Backbone=RN152, Text Encoder=Embedding layer2023.02 | 89.87 | — | — | |
| PolyFormer-LVisual Backbone=Swin-L, Text Encoder=BERT-base2023.02 | 89.77 | — | — | |
| LyricsModel type=Generalist Models, Backbone=Vicuna-13B2023.12 | 89.77 | — | — | |
| UNINEXT-HModel type=Specialist SOTAs (Specialist/Finetuned)2023.06 | 89.63 | — | — | |
| UNINEXT-HModel type=Specialist SOTAs2023.08 | 89.63 | — | — | |
| UNINEXTModel type=Specialist models2023.10 | 89.63 | — | — | |
| UNINEXT-HType=Specialist2023.11 | 89.63 | — | — | |
| UNINEXT-HModel type=Specialist Models2023.12 | 89.63 | — | — | |
| UNINEXT-HType=Specialist2024.09 | 89.6 | — | — | |
| UNINEXT-H2025.12 | 89.6 | — | — | |
| GETok-R1-gridLearning Paradigm=Reinforcement Learning, Evaluation Protocol=Acc@0.52025.12 | 89.6 | 89.9 | — | |
| UNINEXT-HBackbone=Huge2024.12 | 89.6 | — | — | |
| LION-12BEvaluation Setting=Fine-tuning Setting2023.11 | 89.22 | — | — | |
| Qwen2.5-VL-7B2025.12 | 89.1 | — | — | |
| G-DINO-LType=Specialist2024.09 | 89 | — | — | |
| Grounding-DINO-L2025.12 | 89 | — | — | |
| Grounding-DINO-LBackbone=Large2024.12 | 89 | — | — | |
| G-DINO-LModel type=Specialist SOTAs (Specialist/Finetuned)2023.06 | 88.95 | — | — | |
| G-DINO-LModel type=Specialist SOTAs2023.08 | 88.95 | — | — | |
| G-DINO-LModel type=Specialist models2023.10 | 88.95 | — | — | |
| G-DINO-LType=Specialist2023.11 | 88.95 | — | — | |
| G-DINO-LModel type=Specialist Models2023.12 | 88.95 | — | — | |
| Elysium (7B)Resolution=336, Visual token count=1082024.03 | 88.93 | — | — | |
| LION-4BEvaluation Setting=Fine-tuning Setting2023.11 | 88.72 | — | — | |
| VisionReasonerLearning Paradigm=Reinforcement Learning, Evaluation Protocol=Acc@0.52025.12 | 88.7 | 89 | — | |
| MM1.52024.12 | 88.7 | — | — | |
| PolyFormer-BVisual Backbone=Swin-B, Text Encoder=BERT-base2023.02 | 88.6 | — | — | |
| Qwen-VL-7B-ChatModel type=Generalist Models2023.08 | 88.59 | — | — | |
| Qwen-VLModel type=Generalist Models, Backbone=Qwen-7B2023.12 | 88.59 | — | — | |
| Qwen-VL-7BParameters=7B2024.02 | 88.59 | — | — | |
| Qwen-VLType=Generalist2024.09 | 88.3 | — | — | |
| Qwen-VL-7BModel type=Generalist Models2023.08 | 88.25 | — | — | |
| Qwen-VLType=Generalist2023.11 | 88.25 | — | — | |
| Qwen-VL-7BModel Size=7B2024.07 | 88.25 | — | — | |
| GETok-SFTLearning Paradigm=Supervised Fine-Tuning, Evaluation Protocol=Acc@0.52025.12 | 88.2 | 90.9 | — | |
| Ferret-13BEvaluation Setting=Fine-tuning Setting2023.11 | 88.14 | — | — | |
| Ferret-13BType=Generalist2023.11 | 88.14 | — | — | |
| Ferret-13BModel Size=13B2024.07 | 88.14 | — | — | |
| Ferret-13BParameters=13B2024.02 | 88.14 | — | — | |
| InternVL2-8BType=Generalist2024.09 | 87.9 | — | — | |
| InternVL2-8BParameters=8B2024.12 | 87.9 | — | — | |
| GETok-SFT-gridLearning Paradigm=Supervised Fine-Tuning, Evaluation Protocol=Acc@0.52025.12 | 87.8 | 90.8 | — | |
| Shikra-13BModel type=Generalist VL SOTAs (w/o finetuning)2023.06 | 87.79 | — | — | |
| Shikra-13BModel type=Generalist Models2023.08 | 87.79 | — | — | |
| Shikra (13B)Model type=Generalist models2023.10 | 87.79 | — | — | |
| Shikra-13BEvaluation Setting=Fine-tuning Setting2023.11 | 87.79 | — | — | |
| Shikra-13BType=Generalist2023.11 | 87.79 | — | — | |
| Shikra (13B)Resolution=224, Visual token count=2562024.03 | 87.79 | — | — | |
| ShikraModel type=Generalist Models, Backbone=Vicuna-13B2023.12 | 87.79 | — | — | |
| Shikra-13BModel Size=13B2024.07 | 87.79 | — | — | |
| Shikra-13BParameters=13B2024.02 | 87.79 | — | — | |
| PinkEvaluation Setting=Fine-tuning Setting2023.11 | 87.5 | — | — | |
| ShikraType=Generalist2024.09 | 87.4 | — | — | |
| Shikra-7BParameters=7B2024.12 | 87.4 | — | — | |
| Ferret-7BEvaluation Setting=Fine-tuning Setting2023.11 | 87.38 | — | — | |
| Ferret (7B)Resolution=336, Visual token count=6082024.03 | 87.38 | — | — | |
| Ferret-7BModel Size=7B2024.07 | 87.38 | — | — | |
| Shikra-7BModel type=Generalist VL SOTAs (w/o finetuning)2023.06 | 87.36 | — | — | |
| Shikra-7BModel type=Generalist Models2023.08 | 87.36 | — | — | |
| Shikra (7B)Model type=Generalist models2023.10 | 87.36 | — | — | |
| Shikra-7BEvaluation Setting=Fine-tuning Setting2023.11 | 87.36 | — | — | |
| Shikra-7BType=Generalist2023.11 | 87.36 | — | — | |
| Shikra (7B)Resolution=224, Visual token count=2562024.03 | 87.36 | — | — | |
| ShikraModel type=Generalist Models, Backbone=Vicuna-7B2023.12 | 87.36 | — | — | |
| Shikra-7BModel Size=7B2024.07 | 87.36 | — | — | |
| Shikra-7BParameters=7B2024.02 | 87.36 | — | — | |
| GroundingGPT (7B)Resolution=336, Visual token count=5762024.03 | 87.18 | — | — |