Mathematical Reasoning on GaoKao (test)
79.74AccuracyDiPO
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| DiPOModel=Qwen3-4B, Temperature=Constant, Maximum context length=8K tokens2026.01 | 79.74 | 1,490.8 | 59.3 | |
| Qwen2.5-Math-72B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=rm@82024.09 | 76.5 | — | — | |
| BaselineModel=Qwen3-4B, Temperature=Constant, Maximum context length=8K tokens2026.01 | 75.58 | 3,208.1 | 40.2 | |
| Qwen2.5-Math-72B-InstructEvaluation Protocol=TOOL-INTEGRATED REASONING, Decoding Strategy=rm@82024.09 | 75.4 | — | — | |
| Qwen2.5-Math-72B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=maj@82024.09 | 75 | — | — | |
| DiPOModel=Qwen3-8B*, Temperature=Constant, Maximum context length=8K tokens2026.01 | 74.55 | 1,575.2 | 59.5 | |
| Qwen2.5-Math-7B-InstructEvaluation Protocol=TOOL-INTEGRATED REASONING, Decoding Strategy=rm@82024.09 | 72.9 | — | — | |
| Qwen2.5-Math-7B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=rm@82024.09 | 72.2 | — | — | |
| Qwen2.5-Math-72B-InstructEvaluation Protocol=TOOL-INTEGRATED REASONING, Decoding Strategy=maj@82024.09 | 72 | — | — | |
| BaselineModel=Qwen3-8B*, Temperature=Constant, Maximum context length=8K tokens2026.01 | 71.43 | 3,976.4 | 27.9 | |
| Qwen2.5-Math-7B-InstructEvaluation Protocol=TOOL-INTEGRATED REASONING, Decoding Strategy=maj@82024.09 | 70.8 | — | — | |
| Qwen2.5-Math-72B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 68.6 | — | — | |
| Qwen2.5-Math-72B-InstructEvaluation Protocol=TOOL-INTEGRATED REASONING, Decoding Strategy=Greedy2024.09 | 68.5 | — | — | |
| Qwen2.5-Math-1.5B-InstructEvaluation Protocol=TOOL-INTEGRATED REASONING, Decoding Strategy=maj@82024.09 | 68.3 | — | — | |
| Qwen2.5-Math-7B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=maj@82024.09 | 68.1 | — | — | |
| Qwen2-Math-72B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=rm@82024.09 | 67.7 | — | — | |
| GPT-4oCode Integration=No2025.02 | 67.5 | — | — | |
| Qwen2.5-Math-1.5B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=rm@82024.09 | 67.5 | — | — | |
| Qwen2.5-Math-1.5B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=maj@82024.09 | 66.4 | — | — | |
| Qwen2.5-Math-7B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 66.3 | — | — | |
| Qwen2.5-Math-1.5B-InstructEvaluation Protocol=TOOL-INTEGRATED REASONING, Decoding Strategy=rm@82024.09 | 64.1 | — | — | |
| Qwen2.5-Math-7B-InstructEvaluation Protocol=TOOL-INTEGRATED REASONING, Decoding Strategy=Greedy2024.09 | 62.9 | — | — | |
| Qwen2-Math-7B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=rm@82024.09 | 62.7 | — | — | |
| Qwen2.5-Math-1.5B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 62.4 | — | — | |
| OpenAI-o1-previewCode Integration=No2025.02 | 62.1 | — | — | |
| Qwen2-Math-72B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=maj@82024.09 | 61.7 | — | — | |
| Qwen2-Math-72B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 59.8 | — | — | |
| Qwen2.5-Math-1.5B-InstructEvaluation Protocol=TOOL-INTEGRATED REASONING, Decoding Strategy=Greedy2024.09 | 59.6 | — | — | |
| Qwen2-Math-7B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=maj@82024.09 | 59.5 | — | — | |
| Qwen2-Math-1.5B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=rm@82024.09 | 58.2 | — | — | |
| Qwen2-72B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 54.6 | — | — | |
| AutoCode4Math-Qwen2.5Code Integration=Autonomous2025.02 | 51.69 | — | — | |
| DeepSeek-Coder-V2-Lite-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 51.1 | — | — | |
| AutoCode4Math-DeepSeekCode Integration=Autonomous2025.02 | 50.53 | — | — | |
| AutoCode4Math-Qwen2Code Integration=Autonomous2025.02 | 50.13 | — | — | |
| Qwen2-Math-1.5B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=maj@82024.09 | 50.1 | — | — | |
| NuminaMath-72BCode Integration=Yes2025.02 | 49.4 | — | — | |
| Qwen2-Math-7B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 49 | — | — | |
| NuminaMath-7B-CoTCode Integration=No2025.02 | 48.83 | — | — | |
| NuminaMath-72B-CoTEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 47.9 | — | — | |
| Qwen2-Math-1.5B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 46.5 | — | — | |
| Mathstral-7BCode Integration=No2025.02 | 46 | — | — | |
| Dart-Math-DeepSeek-7BCode Integration=No2025.02 | 45.45 | — | — | |
| Qwen-2.5-Base-7BCode Integration=No2025.02 | 45.45 | — | — | |
| DeepseekMath-Instruct-7BCode Integration=Yes2025.02 | 44.68 | — | — | |
| Qwen2Math-Base-7BCode Integration=No2025.02 | 43.37 | — | — | |
| GPT-4o-2024-08-06Evaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 42.6 | — | — | |
| Llama-3.1-70B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 41.7 | — | — | |
| Internlm2-math-plus-mixtral8x7BEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 37.3 | — | — | |
| NuminaMath-7B-CoTEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 36.4 | — | — | |
| Internlm2-math-plus-20BEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 36.1 | — | — | |
| Qwen2-7B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 35.1 | — | — | |
| Dart-Math-Llama3-8BCode Integration=No2025.02 | 34.8 | — | — | |
| Internlm2-math-plus-7BEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 34.5 | — | — | |
| DeepSeekMath-7B-RLEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 33.6 | — | — | |
| TORA-70BCode Integration=Yes2025.02 | 31.7 | — | — | |
| Mathstral-7B-v0.1Evaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 31.6 | — | — | |
| Llama-3.1-8B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 30.4 | — | — | |
| Mammoth-70BCode Integration=Yes2025.02 | 25.2 | — | — | |
| Mammoth-Mistral-7BCode Integration=Yes2025.02 | 22.08 | — | — | |
| Qwen2-1.5B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 17 | — | — |