Mathematical Reasoning on NuminaMath (subset of 5,000 samples)
73.94AccuracyGPT-5
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-5Reasoning Strategy=CoT-only, Tool Integration=None2026.01 | 73.94 | |
| Qwen2.5-VL-7B-VisTIRATraining=Supervised fine-tuned on VisTIRA corpus, Scale=7B2026.01 | 60.97 | |
| Qwen2.5-VL-7B-InstructScale=7B2026.01 | 58.77 |