Reward modeling on EVAL_INSTRUCT 4 steps
2.55Step Completion RateR2VLM
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| R2VLM2026.03 | 2.55 | 45 | |
| Pretrained SPRINT2026.03 | 2.25 | 40 | |
| Step-Completion Based Reward2026.03 | 2.15 | 40 | |
| Qwen2.5-VL-InstructModel Size=7B2026.03 | 1.75 | 25 |
| Method | Links | ||
|---|---|---|---|
| R2VLM2026.03 | 2.55 | 45 | |
| Pretrained SPRINT2026.03 | 2.25 | 40 | |
| Step-Completion Based Reward2026.03 | 2.15 | 40 | |
| Qwen2.5-VL-InstructModel Size=7B2026.03 | 1.75 | 25 |