Reward modeling on EVAL_INSTRUCT 3 steps
2.2Step Completion RateR2VLM
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| R2VLM2026.03 | 2.2 | 60 | |
| Pretrained SPRINT2026.03 | 1.9 | 50 | |
| Step-Completion Based Reward2026.03 | 1.85 | 50 | |
| Qwen2.5-VL-InstructModel Size=7B2026.03 | 1.55 | 35 |
| Method | Links | ||
|---|---|---|---|
| R2VLM2026.03 | 2.2 | 60 | |
| Pretrained SPRINT2026.03 | 1.9 | 50 | |
| Step-Completion Based Reward2026.03 | 1.85 | 50 | |
| Qwen2.5-VL-InstructModel Size=7B2026.03 | 1.55 | 35 |