Mathematical Reasoning on AMC 2023 (test)
81.8AccuracyOpenAI-o1-preview
Evaluation Results
| Method | Links | |
|---|---|---|
| OpenAI-o1-previewCode Integration=No2025.02 | 81.8 | |
| NEMOTRON-NANO-12B-V2 (RLP + Post)Pre-training objective=RLP, Post-training pipeline=SFT + RLVR, Evaluation protocol=Pass@1 (average of 8 runs)2025.09 | 75 | |
| NEMOTRON-NANO-12B-V2 (Base)Pre-training objective=Next-token prediction, Post-training pipeline=None, Evaluation protocol=Pass@1 (average of 8 runs)2025.09 | 70.63 | |
| NEMOTRON-NANO-12B-V2 (Base + Post)Pre-training objective=Next-token prediction, Post-training pipeline=SFT + RLVR, Evaluation protocol=Pass@1 (average of 8 runs)2025.09 | 62.19 | |
| SFTBackbone=Qwen2.5-7B-Instruct2025.03 | 60 | |
| GLOREBackbone=Qwen2.5-7B-Instruct2025.03 | 60 | |
| NEMOTRON-NANO-12B-V2 (RLP)Pre-training objective=RLP, Post-training pipeline=None, Evaluation protocol=Pass@1 (average of 8 runs)2025.09 | 57.19 | |
| RoTBackbone=Qwen2.5-7B-Instruct2025.03 | 52.5 | |
| MathNeuroBackbone=Qwen2.5-7B-Instruct2025.03 | 50 | |
| Zero-shot CoTBackbone=Qwen2.5-7B-Instruct2025.03 | 47.5 | |
| GPT-4oCode Integration=No2025.02 | 45.8 | |
| AutoCode4Math-Qwen2.5Code Integration=Autonomous2025.02 | 45.18 | |
| BoostStepBackbone=Qwen2.5-7B-Instruct2025.03 | 45 | |
| JF-HPOBackbone=Qwen-2.5 7B, Training Dataset=MATH2026.06 | 44.58 | |
| Few-shot CoTBackbone=Qwen2.5-7B-Instruct2025.03 | 42.5 | |
| NuminaMath-72BCode Integration=Yes2025.02 | 40.6 | |
| Qwen-2.5-Base-7BCode Integration=No2025.02 | 39.38 | |
| SFTBackbone=Llama3.1-8B-Instruct2025.03 | 35 | |
| GLOREBackbone=Llama3.1-8B-Instruct2025.03 | 35 | |
| Dart-Math-DeepSeek-7BCode Integration=No2025.02 | 35 | |
| RoTBackbone=Llama3.1-8B-Instruct2025.03 | 32.5 | |
| Zero-shot CoTBackbone=Llama3.1-8B-Instruct2025.03 | 30 | |
| MathNeuroBackbone=Llama3.1-8B-Instruct2025.03 | 30 | |
| AutoCode4Math-Qwen2Code Integration=Autonomous2025.02 | 30 | |
| AutoCode4Math-DeepSeekCode Integration=Autonomous2025.02 | 28.8 | |
| VeRL RecipeBackbone=Qwen-2.5 7B, Training Dataset=MATH2026.06 | 27.71 | |
| Few-shot CoTBackbone=Llama3.1-8B-Instruct2025.03 | 27.5 | |
| BoostStepBackbone=Llama3.1-8B-Instruct2025.03 | 25 | |
| NuminaMath-7B-CoTCode Integration=No2025.02 | 25 | |
| Mammoth-Mistral-7BCode Integration=Yes2025.02 | 20 | |
| Qwen2Math-Base-7BCode Integration=No2025.02 | 19.8 | |
| Dart-Math-Llama3-8BCode Integration=No2025.02 | 17.5 | |
| DeepseekMath-Instruct-7BCode Integration=Yes2025.02 | 17.4 |