Viewpoint Robustness on Progress-Bench 1.0 (test)
-1.1ΔNSEIntern3.5-VL-8B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Intern3.5-VL-8BReasoning Protocol=Training-free Progress Reasoning2026.01 | -1.1 | 10.5 | -0.1 | |
| Intern3.5-VL-4BReasoning Protocol=Training-free Progress Reasoning2026.01 | -0.9 | -21 | -0.5 | |
| GPT-5-miniReasoning Protocol=Training-free Progress Reasoning2026.01 | 1.3 | 5 | 0 | |
| Qwen2.5-VL-3BReasoning Protocol=Training-free Progress Reasoning2026.01 | 4.2 | -14.1 | -3.4 | |
| Intern3.5-VL-14BReasoning Protocol=Training-free Progress Reasoning2026.01 | 4.3 | -64.5 | 0 | |
| ProgressLM-3B-SFTTraining Method=SFT2026.01 | 4.9 | -16.2 | 8.2 | |
| ProgressLM-3B-RLTraining Method=RL2026.01 | 4.9 | -4.7 | 11.6 | |
| GPT-5Reasoning Protocol=Training-free Progress Reasoning2026.01 | 5.9 | 0 | 1.8 | |
| Qwen3-VL-32BReasoning Protocol=Training-free Progress Reasoning2026.01 | 6 | -10.5 | 0.4 | |
| Qwen3-VL-8BReasoning Protocol=Training-free Progress Reasoning2026.01 | 6.1 | -20.4 | 0 | |
| Qwen3-VL-2BReasoning Protocol=Training-free Progress Reasoning2026.01 | 6.2 | -4.1 | -0.1 | |
| Qwen2.5-VL-7BReasoning Protocol=Training-free Progress Reasoning2026.01 | 9.2 | -25.2 | -4.6 | |
| Intern3.5-VL-38BReasoning Protocol=Training-free Progress Reasoning2026.01 | 10.5 | -25.1 | 50.8 | |
| Qwen3-VL-4BReasoning Protocol=Training-free Progress Reasoning2026.01 | 12.5 | -20.7 | 0 | |
| Qwen2.5-VL-72BReasoning Protocol=Training-free Progress Reasoning2026.01 | 14.1 | -33.2 | 13 | |
| Qwen2.5-VL-32BReasoning Protocol=Training-free Progress Reasoning2026.01 | 24.3 | -43.5 | 0 |