Mathematical Reasoning on MGSM-zh (test)
89.6AccuracyJT-Safe-V2-35B
Evaluation Results
| Method | Links | |
|---|---|---|
| JT-Safe-V2-35BParameters=35B2026.05 | 89.6 | |
| SOTA with Equivalent ParametersModel comparison=Equivalent Parameters2026.05 | 89.2 | |
| DeepSeekMath-RLSize=7B, Reasoning Mode=Chain-of-Thought, Source Type=Open-Source, Evaluation Protocol=Top12024.02 | 79.6 | |
| DeepSeekMath-RLSize=7B, Reasoning Mode=Tool-Integrated, Source Type=Open-Source, Evaluation Protocol=Top12024.02 | 78.4 | |
| DeepSeek-LLM-ChatSize=67B, Reasoning Mode=Tool-Integrated, Source Type=Open-Source, Evaluation Protocol=Top12024.02 | 76.4 | |
| DeepSeek-LLM-ChatSize=67B, Reasoning Mode=Chain-of-Thought, Source Type=Open-Source, Evaluation Protocol=Top12024.02 | 74 | |
| DeepSeekMath-InstructSize=7B, Reasoning Mode=Chain-of-Thought, Source Type=Open-Source, Evaluation Protocol=Top12024.02 | 73.2 | |
| DeepSeekMath-InstructSize=7B, Reasoning Mode=Tool-Integrated, Source Type=Open-Source, Evaluation Protocol=Top12024.02 | 72 | |
| MetaMathSize=70B, Reasoning Mode=Chain-of-Thought, Source Type=Open-Source, Evaluation Protocol=Majority Vote (32 candidates)2024.02 | 66.4 | |
| SeaLLM-v2Size=7B, Reasoning Mode=Chain-of-Thought, Source Type=Open-Source, Evaluation Protocol=Majority Vote (32 candidates)2024.02 | 64.8 | |
| WizardMath-v1.0Size=70B, Reasoning Mode=Chain-of-Thought, Source Type=Open-Source, Evaluation Protocol=Majority Vote (32 candidates)2024.02 | 64.8 | |
| ToRASize=34B, Reasoning Mode=Tool-Integrated, Source Type=Open-Source, Evaluation Protocol=Top12024.02 | 41.2 |