Visual Reasoning on LLVIP Infrared (test)
92.3Union ScoreDAPO+Ours
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| DAPO+OursTraining Strategy=Reinforcement Post-Training2026.05 | 92.3 | 89.4 | |
| GRPO+OursTraining Strategy=Reinforcement Post-Training2026.05 | 89.9 | 87.7 | |
| GRPOTraining Strategy=Reinforcement Post-Training2026.05 | 89.7 | 86 | |
| DAPOTraining Strategy=Reinforcement Post-Training2026.05 | 88.8 | 85.5 | |
| Chain-of-ThoughtTraining Strategy=Supervised Post-Training2026.05 | 88.2 | 84.2 | |
| SFTTraining Strategy=Supervised Post-Training2026.05 | 87.1 | 84.5 | |
| VisionReasoner-7BTraining Strategy=Previous Methods2026.05 | 82.5 | 75.2 | |
| Qwen2.5-VL-3BTraining Strategy=Previous Methods2026.05 | 63 | 38.8 | |
| R1-Onevision-7BTraining Strategy=Previous Methods2026.05 | 19.7 | 11.1 | |
| LLaVA-OV-7BTraining Strategy=Previous Methods2026.05 | 4.9 | 2.1 |