Referring Expression Comprehension on VRSBench (test)
51.36Accuracy@0.5RS-HyRe-R1
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| RS-HyRe-R1Base LLM=Qwen2.5-3B, Evaluation Protocol=RL Train2026.04 | 51.36 | 32.25 | |
| GeoChatBase LLM=Vicuna1.5-7B, Evaluation Protocol=Full Amount Fine-tune (36,313 samples)2026.04 | 49.8 | 19.9 | |
| Geo-R1-RECBase LLM=Qwen2.5-3B, Evaluation Protocol=RL Train2026.04 | 49.6 | 29.77 | |
| Geo-R1Base LLM=Qwen2.5-7B, Evaluation Protocol=RL Train2026.04 | 42.03 | 17.62 | |
| GeoReasonBase LLM=Qwen2.5-7B, Evaluation Protocol=RL Train2026.04 | 41.67 | 23.5 | |
| LLaVA-1.5Base LLM=Vicuna1.5-7B, Evaluation Protocol=Full Amount Fine-tune (36,313 samples)2026.04 | 41.6 | 13.6 | |
| Qwen2.5-VL 7BBase LLM=Qwen2.5-7B, Evaluation Protocol=Zero-shot Baseline2026.04 | 41.28 | 23.32 | |
| VLM-R1-RECBase LLM=Qwen2.5-3B, Evaluation Protocol=RL Train2026.04 | 38.77 | 21.18 | |
| Geo-R1-OVDBase LLM=Qwen2.5-3B, Evaluation Protocol=RL Train2026.04 | 38.22 | 20.9 | |
| MiniGPT-v2Base LLM=Vicuna1.5-7B, Evaluation Protocol=Full Amount Fine-tune (36,313 samples)2026.04 | 35.8 | 16.8 | |
| Qwen2.5-VL-SFTBase LLM=Qwen2.5-3B, Evaluation Protocol=RS-task Dataset Fine-tune (1600 samples)2026.04 | 35.39 | 18.58 | |
| R1-VLBase LLM=Qwen2-7B, Evaluation Protocol=RL Train2026.04 | 35.02 | 19.28 | |
| VLM-R1-OVDBase LLM=Qwen2.5-3B, Evaluation Protocol=RL Train2026.04 | 33.51 | 17.68 | |
| Mini-GeminiBase LLM=Gemma-7B, Evaluation Protocol=Full Amount Fine-tune (36,313 samples)2026.04 | 30.1 | 6.8 | |
| TinyRS-R1Base LLM=Qwen2-2B, Evaluation Protocol=RL Train2026.04 | 23.94 | 10.07 | |
| Qwen2.5-VL 3BBase LLM=Qwen2.5-3B, Evaluation Protocol=Zero-shot Baseline2026.04 | 16.89 | 7.36 |