Vision-Intensive Perception on V* Benchmark
84.4Attr ScoreLVR
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| LVRParams=7B2026.01 | 84.4 | — | — | |
| Qwen2.5-VLParams=32B2026.01 | 82.61 | — | — | |
| LaViTParams=3B2026.01 | 82.61 | — | — | |
| ThymeTool Usage=Yes2026.01 | 82.5 | 78.9 | 57.2 | |
| Qwen2.5-VL (Baseline)Params=3B2026.01 | 81.74 | — | — | |
| Naive SFTParams=3B2026.01 | 80.87 | — | — | |
| Ours (SFT)Training Paradigm=SFT, Tool Usage=Yes2026.01 | 80.2 | 78.9 | 55.2 | |
| Ours (RL)Training Paradigm=RL, Tool Usage=Yes2026.01 | 77.8 | 76.3 | 55.1 | |
| R1-onevision-RLTraining Paradigm=RL, Tool Usage=Yes2026.01 | 77.4 | 94.5 | 54.2 | |
| Qwen2.5-VLParams=7B2026.01 | 77.39 | — | — | |
| GPT-4oParams=-2026.01 | 72.5 | — | — | |
| Qwen2.5-VL-7B-InstructModel Scale=7B, Tool Usage=Yes2026.01 | 72.2 | 77.6 | 48.8 | |
| LVR_RLParams=3B2026.01 | 69.6 | — | — | |
| InternVL2.5-8BModel Scale=8B, Tool Usage=Yes2026.01 | 66.1 | 69.7 | 49.6 | |
| LLaVA-OneVision-Qwen2-7BModel Scale=7B, Tool Usage=Yes2026.01 | 60 | 67.1 | 50 | |
| DMLRParams=3B2026.01 | 46.96 | — | — | |
| MM-EurekaQwen-7BModel Scale=7B, Tool Usage=Yes2026.01 | 41.7 | 64.5 | 42.4 | |
| PAPOParams=3B2026.01 | 22.61 | — | — |