Reward modeling on EVAL_INSTRUCT 5 steps
3.38Step Completion RateR2VLM
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| R2VLM2026.03 | 3.38 | 23 | |
| Pretrained SPRINT2026.03 | 3.31 | 23 | |
| Step-Completion Based Reward2026.03 | 2.77 | 23 | |
| Qwen2.5-VL-InstructModel Size=7B2026.03 | 2.61 | 23 |
| Method | Links | ||
|---|---|---|---|
| R2VLM2026.03 | 3.38 | 23 | |
| Pretrained SPRINT2026.03 | 3.31 | 23 | |
| Step-Completion Based Reward2026.03 | 2.77 | 23 | |
| Qwen2.5-VL-InstructModel Size=7B2026.03 | 2.61 | 23 |