Mathematical Reasoning on CN Middle School 24
83.3AccuracyLALP
Evaluation Results
| Method | Links | |
|---|---|---|
| LALPStudent=Qwen2.5-32B-Instruct, Model Configuration=LALP2025.10 | 83.3 | |
| Qwen2.5-Math-7B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=rm@82024.09 | 81.2 | |
| Qwen2.5-Math-1.5B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=rm@82024.09 | 80.2 | |
| Qwen2.5-Math-72B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=rm@82024.09 | 80.2 | |
| Qwen2.5-Math-7B-InstructEvaluation Protocol=TOOL-INTEGRATED REASONING, Decoding Strategy=rm@82024.09 | 80.2 | |
| Qwen2.5-Math-72B-InstructEvaluation Protocol=TOOL-INTEGRATED REASONING, Decoding Strategy=rm@82024.09 | 80.2 | |
| Qwen2-Math-72B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=maj@82024.09 | 79.2 | |
| Qwen2.5-Math-72B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 79.2 | |
| Qwen2.5-Math-72B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=maj@82024.09 | 79.2 | |
| Qwen2-Math-7B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=rm@82024.09 | 78.2 | |
| Qwen2-Math-72B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=rm@82024.09 | 78.2 | |
| Qwen2.5-Math-7B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=maj@82024.09 | 78.2 | |
| Qwen2.5-Math-1.5B-InstructEvaluation Protocol=TOOL-INTEGRATED REASONING, Decoding Strategy=maj@82024.09 | 78.2 | |
| Qwen2.5-Math-1.5B-InstructEvaluation Protocol=TOOL-INTEGRATED REASONING, Decoding Strategy=rm@82024.09 | 78.2 | |
| Qwen2.5-Math-7B-InstructEvaluation Protocol=TOOL-INTEGRATED REASONING, Decoding Strategy=maj@82024.09 | 78.2 | |
| Qwen2.5-Math-72B-InstructEvaluation Protocol=TOOL-INTEGRATED REASONING, Decoding Strategy=Greedy2024.09 | 78.2 | |
| Qwen2.5-Math-72B-InstructEvaluation Protocol=TOOL-INTEGRATED REASONING, Decoding Strategy=maj@82024.09 | 78.2 | |
| Qwen2-Math-72B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 77.2 | |
| Qwen2.5-Math-1.5B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=maj@82024.09 | 77.2 | |
| Qwen2.5-Math-1.5B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 76.2 | |
| NuminaMath-72B-CoTEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 75.2 | |
| Qwen2-Math-1.5B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=rm@82024.09 | 75.2 | |
| Qwen2.5-Math-7B-InstructEvaluation Protocol=TOOL-INTEGRATED REASONING, Decoding Strategy=Greedy2024.09 | 75.2 | |
| Qwen2-72B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 74.3 | |
| Qwen2.5-Math-7B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 73.3 | |
| RandomStudent=Qwen2.5-32B-Instruct, Model Configuration=Random2025.10 | 73.3 | |
| GALPStudent=Qwen2.5-32B-Instruct, Model Configuration=GALP2025.10 | 73.3 | |
| Qwen2-Math-7B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=maj@82024.09 | 72.3 | |
| Qwen2.5-Math-1.5B-InstructEvaluation Protocol=TOOL-INTEGRATED REASONING, Decoding Strategy=Greedy2024.09 | 71.3 | |
| Qwen2-Math-1.5B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=maj@82024.09 | 70.3 | |
| Local LowestStudent=Qwen2.5-32B-Instruct, Model Configuration=Local Lowest2025.10 | 70 | |
| Qwen2-Math-7B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 69.3 | |
| DeepSeekMath-7B-RLEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 67.3 | |
| DeepSeek-Coder-V2-Lite-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 66.3 | |
| Qwen2-Math-1.5B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 66.3 | |
| GPT-4o-2024-08-06Evaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 60.4 | |
| NuminaMath-7B-CoTEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 60.4 | |
| Llama-3.1-70B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 59.4 | |
| Qwen2-7B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 54.5 | |
| Llama-3.1-8B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 43.6 | |
| Mathstral-7B-v0.1Evaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 42.6 | |
| Internlm2-math-plus-mixtral8x7BEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 39.6 | |
| Internlm2-math-plus-20BEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 33.7 | |
| LALPStudent=Qwen2.5-7B-Instruct, Model Configuration=LALP2025.10 | 33.3 | |
| Internlm2-math-plus-7BEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 32.7 | |
| Qwen2-1.5B-InstructEvaluation Protocol=CHAIN-OF-THOUGHT, Decoding Strategy=Greedy2024.09 | 31.7 | |
| RandomStudent=Qwen2.5-7B-Instruct, Model Configuration=Random2025.10 | 23.3 | |
| GALPStudent=Qwen2.5-7B-Instruct, Model Configuration=GALP2025.10 | 23.3 | |
| Local LowestStudent=Qwen2.5-7B-Instruct, Model Configuration=Local Lowest2025.10 | 23.3 | |
| Original ModelStudent=Qwen2.5-32B-Instruct, Model Configuration=Original Model2025.10 | 23.3 | |
| Original ModelStudent=Qwen2.5-7B-Instruct, Model Configuration=Original Model2025.10 | 16.7 |