Referring Expression Comprehension on VRSBench
66.54Unique Accuracy @ IoU=0.5Qwen2.5-VL
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| Qwen2.5-VLBase LLM=Qwen2.5-3B, Evaluation Protocol=Full Amount Fine-tune, Number of Samples=36,313 samples2025.09 | 66.54 | 36.77 | 60.32 | 36.3 | 62.91 | 36.5 | |
| Gemini 3.1 ProBase LLM=Gemini 3.1 Pro, Evaluation Protocol=Zero-shot Baseline2025.09 | 62.44 | 42.72 | 53.92 | 36.5 | 57.55 | 39.15 | |
| Geo-R1 (DAPO)Base LLM=Qwen2.5-3B, Evaluation Protocol=10-shot Fine-tune, Number of Samples=260 samples2025.09 | 59.49 | 37.11 | 47.91 | 28.07 | 52.74 | 31.84 | |
| GeoChatBase LLM=Vicuna1.5-7B, Evaluation Protocol=Full Amount Fine-tune, Number of Samples=36,313 samples2025.09 | 57.4 | 22.6 | 44.5 | 18 | 49.8 | 19.9 | |
| Geo-R1 (GRPO)Base LLM=Qwen2.5-3B, Evaluation Protocol=10-shot Fine-tune, Number of Samples=260 samples2025.09 | 57.27 | 35.61 | 45.81 | 27.03 | 50.59 | 30.61 | |
| Geo-R1 (DAPO)Base LLM=Qwen2.5-3B, Evaluation Protocol=5-shot Fine-tune, Number of Samples=130 samples2025.09 | 55.73 | 32.19 | 44.19 | 24.86 | 49 | 27.92 | |
| Geo-R1 (GRPO)Base LLM=Qwen2.5-3B, Evaluation Protocol=5-shot Fine-tune, Number of Samples=130 samples2025.09 | 54.11 | 31.35 | 42.98 | 23.98 | 47.62 | 27.06 | |
| Geo-R1 (GRPO)Base LLM=Qwen2.5-3B, Evaluation Protocol=1-shot Fine-tune, Number of Samples=26 samples2025.09 | 52.17 | 31.18 | 41.21 | 23.04 | 45.78 | 26.43 | |
| Geo-R1 (DAPO)Base LLM=Qwen2.5-3B, Evaluation Protocol=1-shot Fine-tune, Number of Samples=26 samples2025.09 | 51.72 | 31.68 | 42.13 | 24.5 | 46.13 | 27.5 | |
| LLaVA-1.5Base LLM=Vicuna1.5-7B, Evaluation Protocol=Full Amount Fine-tune, Number of Samples=36,313 samples2025.09 | 51.1 | 16.4 | 34.8 | 11.5 | 41.6 | 13.6 | |
| Qwen2.5-VL w/ thinkingBase LLM=Qwen2.5-3B, Evaluation Protocol=Zero-shot Baseline2025.09 | 46.18 | 26.9 | 35.22 | 18.87 | 39.79 | 22.22 | |
| Qwen2.5-VL w/o thinkingBase LLM=Qwen2.5-3B, Evaluation Protocol=Zero-shot Baseline2025.09 | 43.1 | 25.1 | 33.46 | 18.01 | 37.48 | 20.97 | |
| Qwen2.5-VL-SFTBase LLM=Qwen2.5-3B, Evaluation Protocol=10-shot Fine-tune, Number of Samples=260 samples2025.09 | 41.81 | 18.59 | 35.78 | 17.2 | 38.29 | 17.78 | |
| Mini-GeminiBase LLM=Gemma-7B, Evaluation Protocol=Full Amount Fine-tune, Number of Samples=36,313 samples2025.09 | 41.1 | 9.6 | 22.3 | 4.9 | 30.1 | 6.8 | |
| MiniGPT-v2Base LLM=Vicuna1.5-7B, Evaluation Protocol=Full Amount Fine-tune, Number of Samples=36,313 samples2025.09 | 40.7 | 18.9 | 32.4 | 15.2 | 35.8 | 16.8 | |
| Qwen2.5-VL-SFTBase LLM=Qwen2.5-3B, Evaluation Protocol=5-shot Fine-tune, Number of Samples=130 samples2025.09 | 36.98 | 16.61 | 33.94 | 17.17 | 35.21 | 16.94 | |
| Qwen2.5-VL-SFTBase LLM=Qwen2.5-3B, Evaluation Protocol=1-shot Fine-tune, Number of Samples=26 samples2025.09 | 34.32 | 18.87 | 31.62 | 16.35 | 32.75 | 17.4 | |
| GPT-4VBase LLM=GPT-4, Evaluation Protocol=Zero-shot Baseline2025.09 | 8.6 | 2.2 | 2.5 | 0.4 | 5.1 | 1.1 |