Progress Reasoning on Progress-Bench Cross-View 1.0 (test)
15.2NSEProgressLM-3B-RL
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| ProgressLM-3B-RLTraining Method=RL2026.01 | 15.2 | 88.8 | 11.7 | |
| ProgressLM-3B-SFTTraining Method=SFT2026.01 | 20.4 | 67.8 | 8.8 | |
| GPT-5Reasoning Protocol=Training-free Progress Reasoning2026.01 | 20.5 | 89.4 | 1.8 | |
| GPT-5-miniReasoning Protocol=Training-free Progress Reasoning2026.01 | 21.1 | 89.1 | 0.4 | |
| Qwen3-VL-32BReasoning Protocol=Training-free Progress Reasoning2026.01 | 21.9 | 77.8 | 0.4 | |
| Qwen3-VL-8BReasoning Protocol=Training-free Progress Reasoning2026.01 | 25.2 | 61.3 | 0 | |
| Qwen2.5-VL-72BReasoning Protocol=Training-free Progress Reasoning2026.01 | 30.9 | 50.7 | 13.4 | |
| Qwen2.5-VL-3BReasoning Protocol=Training-free Progress Reasoning2026.01 | 33.4 | 28.9 | 6.5 | |
| Qwen3-VL-4BReasoning Protocol=Training-free Progress Reasoning2026.01 | 34.1 | 51.6 | 0 | |
| Qwen2.5-VL-7BReasoning Protocol=Training-free Progress Reasoning2026.01 | 36.5 | 26.7 | 26 | |
| Intern3.5-VL-38BReasoning Protocol=Training-free Progress Reasoning2026.01 | 38.1 | 49.7 | 51.3 | |
| Intern3.5-VL-4BReasoning Protocol=Training-free Progress Reasoning2026.01 | 43.5 | 12.1 | 0.2 | |
| Qwen2.5-VL-32BReasoning Protocol=Training-free Progress Reasoning2026.01 | 45.7 | 29.4 | 0 | |
| Intern3.5-VL-8BReasoning Protocol=Training-free Progress Reasoning2026.01 | 64.1 | 15.7 | 0.4 | |
| Qwen3-VL-2BReasoning Protocol=Training-free Progress Reasoning2026.01 | 66.2 | 30.9 | 0 | |
| Intern3.5-VL-14BReasoning Protocol=Training-free Progress Reasoning2026.01 | 66.4 | -40.3 | 0 |