Mathematical Reasoning on WeMath 525 samples
78.5AccuracyDoubao-Seed-1.6
Evaluation Results
| Method | Links | |
|---|---|---|
| Doubao-Seed-1.6Model Category=Closed-Source SOTA Models2026.05 | 78.5 | |
| Qwen3-VL-235B-ThinkingModel Category=Open-Source Large Baselines2026.05 | 74.9 | |
| Gemini3-proModel Category=Closed-Source SOTA Models2026.05 | 73.8 | |
| GPT-5-20250807Model Category=Closed-Source SOTA Models2026.05 | 71 | |
| GLM-4.5VModel Category=Closed-Source SOTA Models2026.05 | 68.8 | |
| OSTBase Model=Qwen3-VL-8B-Instruct, Sampling Ratio=Best-20%2026.05 | 59 | |
| DEITABase Model=Qwen3-VL-8B-Instruct, Sampling Ratio=Top-20%2026.05 | 58.6 | |
| Qwen3-VL-8B-Instruct + LLM-as-a-JudgeSampling Ratio=20%2026.05 | 58.1 | |
| Qwen3-VL-8B-InstructStrategy=Base2026.05 | 57.5 | |
| Qwen3-VL-8B-Instruct + RandomSampling Ratio=20%2026.05 | 56.8 | |
| OSTBase Model=Qwen3-VL-4B-Instruct, Sampling Ratio=Best-20%2026.05 | 56 | |
| DEITABase Model=Qwen3-VL-4B-Instruct, Sampling Ratio=Top-20%2026.05 | 55.2 | |
| Qwen3-VL-8B-Instruct + Full SFTSampling Ratio=100%2026.05 | 55.2 | |
| Qwen3-VL-4B-Instruct + LLM-as-a-JudgeSampling Ratio=20%2026.05 | 55 | |
| Qwen3-VL-4B-InstructStrategy=Base2026.05 | 54.3 | |
| Qwen3-VL-4B-Instruct + RandomSampling Ratio=20%2026.05 | 53.9 | |
| Qwen3-VL-4B-Instruct + Full SFTSampling Ratio=100%2026.05 | 52.6 | |
| Kimi-vl-A3B-thinkingModel Category=Open-Source Large Baselines2026.05 | 47 | |
| OSTBase Model=Qwen3-VL-2B-Instruct, Sampling Ratio=Best-20%2026.05 | 37.5 | |
| DEITABase Model=Qwen3-VL-2B-Instruct, Sampling Ratio=Top-20%2026.05 | 36.5 | |
| Qwen3-VL-2B-Instruct + LLM-as-a-JudgeSampling Ratio=20%2026.05 | 36.2 | |
| Qwen3-VL-2B-InstructStrategy=Base2026.05 | 35.8 | |
| Qwen3-VL-2B-Instruct + RandomSampling Ratio=20%2026.05 | 35.4 | |
| Qwen3-VL-2B-Instruct + Full SFTSampling Ratio=100%2026.05 | 34.9 |