Mathematical Reasoning on OlympiadBench OE_TO_maths_en_COMP
74.04AccuracyQwen3-235B-A22B
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3-235B-A22BModel Category=Large Language Models, Model Scale=235B-A22B, Reasoning Strategy=Thinking2025.12 | 74.04 | |
| FIGR2025.12 | 72.4 | |
| Qwen3-32BModel Category=Large Language Models, Model Scale=32B, Reasoning Strategy=Thinking2025.12 | 71.96 | |
| Text-only RLMode=Text-only, Reasoning Strategy=RL2025.12 | 70.33 | |
| GLM-4.5VModel Category=Large Vision-Language Models, Model Scale=108B2025.12 | 70.03 | |
| Qwen3-VL-32B-InstructModel Category=Large Vision-Language Models, Model Scale=32B, Reasoning Strategy=Instruct2025.12 | 69.73 | |
| Qwen3-VL-8B-InstructModel Category=Large Vision-Language Models, Model Scale=8B, Reasoning Strategy=Instruct2025.12 | 67.21 | |
| Qwen3-32BModel Category=Large Language Models, Model Scale=32B, Reasoning Strategy=Non-Thinking2025.12 | 52.08 | |
| Bagel-7B-MoTModel Category=Unified Multimodal Models, Model Scale=7B, Reasoning Strategy=MoT2025.12 | 26.11 | |
| DeepEyesModel Category=Tool-Augmented Vision-Language Models2025.12 | 13.65 | |
| Bagel-Zebra-CoTModel Category=Unified Multimodal Models, Model Scale=7B, Reasoning Strategy=CoT2025.12 | 8.46 | |
| Chain-of-FocusModel Category=Tool-Augmented Vision-Language Models2025.12 | 0.45 |