Mathematical Problem Solving on MinervaMath (test)
38.5AccuracyConfidence + AXON
Evaluation Results
| Method | Links | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Confidence + AXONBackbone=Dream-v0-Instruct-7B, Decoding strategy=Confidence + AXON2026.06 | 38.5 | — | — | — | — | 24.04 | 112.1 | — | — | — | — | — | — | |
| DAWN + AXONBackbone=Dream-v0-Instruct-7B, Decoding strategy=DAWN + AXON2026.06 | 38.28 | — | — | — | — | 28.89 | 93 | — | — | — | — | — | — | |
| ConfidenceBackbone=Dream-v0-Instruct-7B, Decoding strategy=Confidence2026.06 | 38.24 | — | — | — | — | 24.66 | 114.8 | — | — | — | — | — | — | |
| DAWNBackbone=Dream-v0-Instruct-7B, Decoding strategy=DAWN2026.06 | 38.22 | — | — | — | — | 28.65 | 94.5 | — | — | — | — | — | — | |
| LocalLeapBackbone=Dream-v0-Instruct-7B, Decoding strategy=LocalLeap2026.06 | 38.1 | — | — | — | — | 28.22 | 97.7 | — | — | — | — | — | — | |
| LocalLeap + AXONBackbone=Dream-v0-Instruct-7B, Decoding strategy=LocalLeap + AXON2026.06 | 37.9 | — | — | — | — | 28.38 | 91.2 | — | — | — | — | — | — | |
| OriginalBackbone=Dream-v0-Instruct-7B, Decoding strategy=Original2026.06 | 37.8 | — | — | — | — | 11.5 | 256 | — | — | — | — | — | — | |
| DAWN + AXONCVRBackbone=Dream-v0-Instruct-7B, Decoding strategy=DAWN + AXONCVR2026.06 | 37.76 | — | — | — | — | 28.86 | 91.2 | — | — | — | — | — | — | |
| Confidence + AXONCVRBackbone=Dream-v0-Instruct-7B, Decoding strategy=Confidence + AXONCVR2026.06 | 37.74 | — | — | — | — | 24.2 | 110 | — | — | — | — | — | — | |
| LocalLeap + AXONCVRBackbone=Dream-v0-Instruct-7B, Decoding strategy=LocalLeap + AXONCVR2026.06 | 37.42 | — | — | — | — | 28.32 | 89.5 | — | — | — | — | — | — | |
| ConfidenceBackbone=LLaDA-1.5, Decoding strategy=Confidence2026.06 | 33.38 | — | — | — | — | 22.31 | 95.4 | — | — | — | — | — | — | |
| OriginalBackbone=LLaDA-1.5, Decoding strategy=Original2026.06 | 33.36 | — | — | — | — | 8.35 | 256 | — | — | — | — | — | — | |
| Confidence + AXONBackbone=LLaDA-1.5, Decoding strategy=Confidence + AXON2026.06 | 33.24 | — | — | — | — | 22.61 | 92.7 | — | — | — | — | — | — | |
| OriginalBackbone=LLaDA-8B-Instruct, Decoding strategy=Original2026.06 | 33.2 | — | — | — | — | 9.51 | 256 | — | — | — | — | — | — | |
| Confidence + AXONCVRBackbone=LLaDA-1.5, Decoding strategy=Confidence + AXONCVR2026.06 | 33.02 | — | — | — | — | 23.06 | 91.1 | — | — | — | — | — | — | |
| ConfidenceBackbone=LLaDA-8B-Instruct, Decoding strategy=Confidence2026.06 | 32.98 | — | — | — | — | 25.14 | 96.9 | — | — | — | — | — | — | |
| LocalLeap + AXONBackbone=LLaDA-1.5, Decoding strategy=LocalLeap + AXON2026.06 | 32.88 | — | — | — | — | 25.86 | 79.7 | — | — | — | — | — | — | |
| Confidence + AXONBackbone=LLaDA-8B-Instruct, Decoding strategy=Confidence + AXON2026.06 | 32.76 | — | — | — | — | 25.75 | 92.7 | — | — | — | — | — | — | |
| LocalLeap + AXONBackbone=LLaDA-8B-Instruct, Decoding strategy=LocalLeap + AXON2026.06 | 32.62 | — | — | — | — | 29.75 | 79.1 | — | — | — | — | — | — | |
| LocalLeapBackbone=LLaDA-8B-Instruct, Decoding strategy=LocalLeap2026.06 | 32.42 | — | — | — | — | 31.61 | 75.6 | — | — | — | — | — | — | |
| LocalLeapBackbone=LLaDA-1.5, Decoding strategy=LocalLeap2026.06 | 32.4 | — | — | — | — | 27.64 | 75.7 | — | — | — | — | — | — | |
| DAWNBackbone=LLaDA-8B-Instruct, Decoding strategy=DAWN2026.06 | 32.36 | — | — | — | — | 32.83 | 72.5 | — | — | — | — | — | — | |
| LocalLeap + AXONCVRBackbone=LLaDA-1.5, Decoding strategy=LocalLeap + AXONCVR2026.06 | 32.34 | — | — | — | — | 26.47 | 78.1 | — | — | — | — | — | — | |
| Confidence + AXONCVRBackbone=LLaDA-8B-Instruct, Decoding strategy=Confidence + AXONCVR2026.06 | 32.3 | — | — | — | — | 26.23 | 91 | — | — | — | — | — | — | |
| LocalLeap + AXONCVRBackbone=LLaDA-8B-Instruct, Decoding strategy=LocalLeap + AXONCVR2026.06 | 32.16 | — | — | — | — | 30.52 | 77.2 | — | — | — | — | — | — | |
| DAWN + AXONBackbone=LLaDA-1.5, Decoding strategy=DAWN + AXON2026.06 | 32 | — | — | — | — | 30.64 | 68 | — | — | — | — | — | — | |
| DAWN + AXONBackbone=LLaDA-8B-Instruct, Decoding strategy=DAWN + AXON2026.06 | 31.84 | — | — | — | — | 35.36 | 67 | — | — | — | — | — | — | |
| DAWNBackbone=LLaDA-1.5, Decoding strategy=DAWN2026.06 | 31.8 | — | — | — | — | 28.63 | 72.9 | — | — | — | — | — | — | |
| DAWN + AXONCVRBackbone=LLaDA-1.5, Decoding strategy=DAWN + AXONCVR2026.06 | 31.3 | — | — | — | — | 31.6 | 66 | — | — | — | — | — | — | |
| DAWN + AXONCVRBackbone=LLaDA-8B-Instruct, Decoding strategy=DAWN + AXONCVR2026.06 | 30.84 | — | — | — | — | 36.57 | 64.7 | — | — | — | — | — | — | |
| PASERPruning Scheme=LLM-Pruner (25%)2025.02 | 21.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PASERPruning Scheme=Wanda (2:4)2025.02 | 20.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PASERPruning Scheme=SliceGPT (25%)2025.02 | 20.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PASERPruning Scheme=SparseGPT (50%)2025.02 | 20.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| NuggetsPruning Scheme=LLM-Pruner (25%)2025.02 | 19.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| IFDPruning Scheme=LLM-Pruner (25%)2025.02 | 19.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Full DataPruning Scheme=LLM-Pruner (25%)2025.02 | 19.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| NuggetsPruning Scheme=Wanda (2:4)2025.02 | 19 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Instruction MiningPruning Scheme=LLM-Pruner (25%)2025.02 | 18.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| IFDPruning Scheme=Wanda (2:4)2025.02 | 18.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| NuggetsPruning Scheme=SparseGPT (50%)2025.02 | 18.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| NuggetsPruning Scheme=SliceGPT (25%)2025.02 | 18.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Full DataPruning Scheme=Wanda (2:4)2025.02 | 18.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| IFDPruning Scheme=SparseGPT (50%)2025.02 | 18.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| IFDPruning Scheme=SliceGPT (25%)2025.02 | 18.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Instruction MiningPruning Scheme=Wanda (2:4)2025.02 | 18.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Full DataPruning Scheme=SparseGPT (50%)2025.02 | 18.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| RandomPruning Scheme=LLM-Pruner (25%)2025.02 | 18.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Full DataPruning Scheme=SliceGPT (25%)2025.02 | 18.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Instruction MiningPruning Scheme=SparseGPT (50%)2025.02 | 18.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Instruction MiningPruning Scheme=SliceGPT (25%)2025.02 | 18.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| RandomPruning Scheme=Wanda (2:4)2025.02 | 18.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| RandomPruning Scheme=SparseGPT (50%)2025.02 | 17.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| w/o TrainingPruning Scheme=LLM-Pruner (25%)2025.02 | 17.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| RandomPruning Scheme=SliceGPT (25%)2025.02 | 17.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| w/o TrainingPruning Scheme=Wanda (2:4)2025.02 | 17.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| w/o TrainingPruning Scheme=SparseGPT (50%)2025.02 | 17.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| w/o TrainingPruning Scheme=SliceGPT (25%)2025.02 | 16.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Base (No RL)Model=Llama3.2-3B-Instruct, Verifier Setting=None2026.02 | — | — | — | 4.8 | — | — | — | — | — | — | — | — | — | |
| Base (No RL)Model=Qwen2.5-Math-7B, Verifier Setting=None2026.02 | — | — | — | 11.8 | — | — | — | — | — | — | — | — | — | |
| Base ModelBase Model=Qwen2.5-Math-7B, Temperature=0.6, top-p=0.95, Max tokens=4096, Scoring=exact any-boxed answer scoring2026.06 | — | 57.6 | 24.4 | 54.5 | 61.9 | — | — | 34.8 | 45.3 | 67.9 | 72.6 | 76.6 | 80.9 | |
| BBGBase Model=Qwen2.5-Math-7B, Temperature=0.6, top-p=0.95, Max tokens=4096, Scoring=exact any-boxed answer scoring2026.06 | — | 62.1 | 37.1 | 58.8 | 64.5 | — | — | 45.3 | 52.4 | 69.2 | 73.3 | 77 | 81.1 | |
| GRPOBase Model=Qwen2.5-Math-7B, Temperature=0.6, top-p=0.95, Max tokens=4096, Scoring=exact any-boxed answer scoring2026.06 | — | 58.7 | 39.2 | 55.1 | 59.9 | — | — | 45.2 | 50.4 | 64.2 | 68 | 71.6 | 74.6 | |
| No CurriculumModel=Llama3.2-3B-Instruct, Verifier Setting=Majority Vote (w/o. Verifier)2026.02 | — | — | — | 10.7 | — | — | — | — | — | — | — | — | — | |
| No CurriculumModel=Llama3.2-3B-Instruct, Verifier Setting=Entropy (w/o. Verifier)2026.02 | — | — | — | 2.2 | — | — | — | — | — | — | — | — | — | |
| No CurriculumModel=Qwen2.5-Math-7B, Verifier Setting=Majority Vote (w/o. Verifier)2026.02 | — | — | — | 23.2 | — | — | — | — | — | — | — | — | — | |
| No CurriculumModel=Qwen2.5-Math-7B, Verifier Setting=Entropy (w/o. Verifier)2026.02 | — | — | — | 25.4 | — | — | — | — | — | — | — | — | — | |
| PRM (Process-Aware)Base Policy=Qwen2.5-Math-7B2025.12 | — | 35.1 | 37.1 | 46.81 | 48.01 | — | — | — | — | — | — | — | — | |
| PRM-CoT (Process-Aware)Base Policy=Qwen2.5-Math-7B2025.12 | — | 37.33 | 40.1 | 49.07 | 52.9 | — | — | — | — | — | — | — | — | |
| REINFORCEBase Model=Qwen2.5-Math-7B, Temperature=0.6, top-p=0.95, Max tokens=4096, Scoring=exact any-boxed answer scoring2026.06 | — | 57.9 | 37.3 | 55.7 | 60.1 | — | — | 44.2 | 50.5 | 63.9 | 67.2 | 70 | 72.4 | |
| RLVR (Ground Truth)Base Policy=Qwen2.5-Math-7B2025.12 | — | 35.15 | 37.5 | 49.9 | 52.71 | — | — | — | — | — | — | — | — | |
| SFT (Baseline)Base Policy=Qwen2.5-Math-7B2025.12 | — | 33.58 | 34.2 | 43.03 | 46.6 | — | — | — | — | — | — | — | — | |
| VI-CuRLModel=Llama3.2-3B-Instruct, Verifier Setting=Majority Vote (w/o. Verifier)2026.02 | — | — | — | 14 | — | — | — | — | — | — | — | — | — | |
| VI-CuRLModel=Llama3.2-3B-Instruct, Verifier Setting=Entropy (w/o. Verifier)2026.02 | — | — | — | 12.9 | — | — | — | — | — | — | — | — | — | |
| VI-CuRLModel=Qwen2.5-Math-7B, Verifier Setting=Majority Vote (w/o. Verifier)2026.02 | — | — | — | 27.9 | — | — | — | — | — | — | — | — | — | |
| VI-CuRLModel=Qwen2.5-Math-7B, Verifier Setting=Entropy (w/o. Verifier)2026.02 | — | — | — | 19.5 | — | — | — | — | — | — | — | — | — | |
| W-RFBase Model=Qwen2.5-Math-7B, Temperature=0.6, top-p=0.95, Max tokens=4096, Scoring=exact any-boxed answer scoring2026.06 | — | 60.7 | 39.4 | 57.6 | 62.7 | — | — | 46.2 | 52.1 | 67 | 70.6 | 74 | 76.8 |