Point Grounding on PACO (test)
61Patch AccuracyQwen2.5-VL-7B Ours
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen2.5-VL-7B OursModel Category=MLLM w/ pointing ability, Pointing Method=Ours2026.06 | 61 | 47.9 | |
| Qwen2.5-VL-7B (Ours)Model Category=MLLM w/ pointing ability, Evaluation Mode=Ours (Q-Synth + A2P)2026.06 | 61 | 47.9 | |
| Molmo-7B OursModel Category=MLLM w/ pointing ability, Pointing Method=Ours2026.06 | 60.3 | 51 | |
| Molmo-7B (Ours)Model Category=MLLM w/ pointing ability, Evaluation Mode=Ours (Q-Synth + A2P)2026.06 | 60.3 | 51 | |
| Molmo-7B text pointingModel Category=MLLM w/ pointing ability, Pointing Method=text pointing2026.06 | 55.9 | 48.7 | |
| Molmo-7B text pointingModel Category=MLLM w/ pointing ability, Evaluation Mode=text pointing2026.06 | 55.9 | 48.7 | |
| First-Gen-MLLM OursModel Category=MLLM w/o pointing ability, Pointing Method=Ours2026.06 | 54.4 | 46.3 | |
| LLaVA-1.5-7B (Ours)Model Category=MLLM w/o pointing ability, Evaluation Mode=Ours (Q-Synth + A2P)2026.06 | 54.4 | 46.3 | |
| GLaMM-7BModel Category=Segmentation-based Models2026.06 | 54 | 45.6 | |
| Molmo-7B attention pointingModel Category=MLLM w/ pointing ability, Pointing Method=attention pointing2026.06 | 51.7 | 42.8 | |
| Molmo-7B attention pointingModel Category=MLLM w/ pointing ability, Evaluation Mode=attention pointing2026.06 | 51.7 | 42.8 | |
| Qwen2.5-VL-7B text pointingModel Category=MLLM w/ pointing ability, Pointing Method=text pointing2026.06 | 49.1 | 40.7 | |
| Qwen2.5-VL-7B text pointingModel Category=MLLM w/ pointing ability, Evaluation Mode=text pointing2026.06 | 49.1 | 40.7 | |
| Qwen2.5-VL-7B attention pointingModel Category=MLLM w/ pointing ability, Pointing Method=attention pointing2026.06 | 42.4 | 30.9 | |
| Qwen2.5-VL-7B attention pointingModel Category=MLLM w/ pointing ability, Evaluation Mode=attention pointing2026.06 | 42.4 | 30.9 | |
| VLPartModel Category=Segmentation-based Models2026.06 | 41.9 | 38.1 | |
| VLPartModel Category=Segmentation-based Models2026.06 | 41.9 | 38.1 | |
| LISA-7BModel Category=Segmentation-based Models2026.06 | 34.5 | 29 | |
| First-Gen-MLLM attention pointingModel Category=MLLM w/o pointing ability, Pointing Method=attention pointing2026.06 | 23 | 18.3 | |
| LLaVA-1.5-7B attention pointingModel Category=MLLM w/o pointing ability, Evaluation Mode=attention pointing2026.06 | 23 | 18.3 | |
| First-Gen-MLLM text pointingModel Category=MLLM w/o pointing ability, Pointing Method=text pointing2026.06 | 8.5 | 6.8 | |
| LLaVA-1.5-7B text pointingModel Category=MLLM w/o pointing ability, Evaluation Mode=text pointing2026.06 | 8.5 | 6.8 | |
| X-DecoderModel Category=Segmentation-based Models2026.06 | 3.1 | 2.5 | |
| X-DecoderModel Category=Segmentation-based Models2026.06 | 3.1 | 2.5 |