Fine-Grained Perception & Understanding on LISA-Grounding (test)
74.79AccuracyQwen2.5-VL-7B + Visual Jigsaw
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2.5-VL-7B + Visual JigsawModel Scale=7B, Training Strategy=+ Visual Jigsaw2026.04 | 74.79 | |
| Qwen2.5-VL-7B + SSL-R1Model Scale=7B, Training Strategy=+ SSL-R12026.04 | 74.07 | |
| ThinkLite-VL-7BModel Scale=7B, Training Strategy=Representative Reasoning Model2026.04 | 73.7 | |
| VL-Cogito-7BModel Scale=7B, Training Strategy=Representative Reasoning Model2026.04 | 72.26 | |
| Qwen2.5-VL-7BModel Scale=7B, Training Strategy=Baseline2026.04 | 71.89 | |
| LLaVA-Critic-R1-7BModel Scale=7B, Training Strategy=Representative Reasoning Model2026.04 | 68.52 | |
| Qwen2.5-VL-3B + Visual JigsawModel Scale=3B, Training Strategy=+ Visual Jigsaw2026.04 | 63.03 | |
| Qwen2.5-VL-3B + SSL-R1Model Scale=3B, Training Strategy=+ SSL-R12026.04 | 61.58 | |
| Qwen2.5-VL-3BModel Scale=3B, Training Strategy=Baseline2026.04 | 59.59 |