Multidisciplinary Reasoning on MMMU Pro 4
62.03Accuracy@1Qwen3-VL-8B-Instruct (Teacher)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen3-VL-8B-Instruct (Teacher)Model Size=8B, Decoding=greedy2026.05 | 62.03 | — | |
| VGS On-Policy DistillationStudent Model=Qwen3-VL-4B-Instruct, Teacher Model=Qwen3-VL-8B-Instruct2026.05 | 56.86 | 56.86 | |
| Standard On-Policy DistillationStudent Model=Qwen3-VL-4B-Instruct, Teacher Model=Qwen3-VL-8B-Instruct2026.05 | 55.79 | 56.43 | |
| Qwen3-VL-4B-Instruct (Initial Student)Model Size=4B, Decoding=greedy2026.05 | 48.93 | — | |
| VGS On-Policy DistillationStudent Model=Qwen3-VL-2B-Instruct, Teacher Model=Qwen3-VL-8B-Instruct2026.05 | 48.07 | 48.34 | |
| Standard On-Policy DistillationStudent Model=Qwen3-VL-2B-Instruct, Teacher Model=Qwen3-VL-8B-Instruct2026.05 | 45.83 | 47.33 | |
| Qwen3-VL-2B-Instruct (Initial Student)Model Size=2B, Decoding=greedy2026.05 | 34.51 | — |