In-the-wild model generalization on Human Bench Average
57.9NSE ScoreQwen3VL-2B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Qwen3VL-2BModel Scale=2B2026.01 | 57.9 | 19.6 | 0.04 | |
| Intern3.5-VL-8BModel Scale=8B2026.01 | 47.2 | 35.7 | 1.4 | |
| Qwen2.5VL-3BModel Scale=3B2026.01 | 40.3 | 32.6 | 11.9 | |
| Intern3.5-VL-4BModel Scale=4B2026.01 | 34.2 | 42.8 | 0.5 | |
| Intern3.5-VL-38BModel Scale=38B2026.01 | 33.3 | 52.1 | 10.9 | |
| Intern3.5-VL-14BModel Scale=14B2026.01 | 32.2 | 53 | 4.3 | |
| Qwen2.5VL-32BModel Scale=32B2026.01 | 29.2 | 60.2 | 0.1 | |
| Qwen2.5VL-7BModel Scale=7B2026.01 | 29.2 | 55.4 | 6 | |
| ProgressLM-SFT-3BModel Scale=3B, Training Protocol=SFT2026.01 | 26 | 63 | 4.5 | |
| Qwen3VL-4BModel Scale=4B2026.01 | 24 | 74.4 | 0 | |
| Qwen2.5VL-72BModel Scale=72B2026.01 | 23.7 | 76.9 | 1.1 | |
| ProgressLM-RL-3BModel Scale=3B, Training Protocol=RL2026.01 | 23.2 | 67.5 | 6.1 | |
| Qwen3VL-8BModel Scale=8B2026.01 | 22.3 | 75.1 | 0.04 | |
| Qwen3VL-32BModel Scale=32B2026.01 | 21.5 | 84.8 | 0.04 |