Accuracy on MATH500
96.2AccuracyEXAONE-Deep-32B
Evaluation Results
| Method | Links | |
|---|---|---|
| EXAONE-Deep-32BSampling=Bo∞ (Best-of-infinity sampling)2025.09 | 96.2 | |
| GPT-OSS-20BSampling=Bo∞ (Best-of-infinity sampling), reasoning mode=medium2025.09 | 96 | |
| Qwen3-30B-A3B-Thinking-2507Sampling=Bo∞ (Best-of-infinity sampling)2025.09 | 96 | |
| NVIDIA-Nemotron-Nano-9B-v2Sampling=Bo∞ (Best-of-infinity sampling)2025.09 | 95.6 | |
| Qwen3-14BSampling=Bo∞ (Best-of-infinity sampling)2025.09 | 95.6 | |
| Qwen3-30B-A3B-Thinking-2507Sampling=Bo1 (Best-of-1 sampling)2025.09 | 95.4 | |
| MetaStone-S1-32BSampling=Bo∞ (Best-of-infinity sampling)2025.09 | 95 | |
| MetaStone-S1-32BSampling=Bo1 (Best-of-1 sampling)2025.09 | 94.7 | |
| Qwen3-14BSampling=Bo1 (Best-of-1 sampling)2025.09 | 94.6 | |
| EXAONE-Deep-32BSampling=Bo1 (Best-of-1 sampling)2025.09 | 94.5 | |
| Phi-4-reasoningSampling=Bo∞ (Best-of-infinity sampling)2025.09 | 94.4 | |
| NVIDIA-Nemotron-Nano-9B-v2Sampling=Bo1 (Best-of-1 sampling)2025.09 | 93.8 | |
| GPT-OSS-20BSampling=Bo1 (Best-of-1 sampling), reasoning mode=medium2025.09 | 92.8 | |
| DVPOTraining Domain=Math Domain2025.12 | 90 | |
| Robust BellmanTraining Domain=Math Domain2025.12 | 89.4 | |
| GRPOTraining Domain=Math Domain2025.12 | 89.2 | |
| Reinforce++Training Domain=Math Domain2025.12 | 89 | |
| Dr.GRPOTraining Domain=Math Domain2025.12 | 87.8 | |
| Phi-4-reasoningSampling=Bo1 (Best-of-1 sampling)2025.09 | 87.8 | |
| BaseTraining Domain=Math Domain2025.12 | 87.4 | |
| PPOTraining Domain=Math Domain2025.12 | 86.4 | |
| Eurus-2-7B-PRIME-R-TAPtraining=R-TAP integrated2026.03 | 83.5 | |
| INTUITORBackbone=Qwen3-14B, Inference Template=chat inference template2025.05 | 83.4 | |
| Qwen2.5-Math-7B-Inst.2026.03 | 79.8 | |
| Qwen3-14BBackbone=Qwen3-14B, Inference Template=chat inference template2025.05 | 79.4 | |
| Eurus-2-7B-PRIME2026.03 | 78.2 | |
| INTUITORBackbone=Qwen2.5-14B, Inference Template=chat inference template2025.05 | 77 | |
| CAPO (GRPO)Backbone=Qwen2.5-7B-Math2025.12 | 76.8 | |
| GPT-4o2026.03 | 76.4 | |
| GRPOBackbone=Qwen2.5-14B, Inference Template=chat inference template2025.05 | 75.8 | |
| GRPOBackbone=Qwen2.5-7B-Math2025.12 | 75.2 | |
| GRPOBackbone=Qwen2.5-7B, Inference Template=chat inference template2025.05 | 75 | |
| INTUITORBackbone=Qwen2.5-7B, Inference Template=chat inference template2025.05 | 75 | |
| CAPO (RLOO)Backbone=Qwen2.5-7B-Math2025.12 | 74.8 | |
| RLOOBackbone=Qwen2.5-7B-Math2025.12 | 73.8 | |
| RLOO2026.03 | 73.2 | |
| CAPO (PPO)Backbone=Qwen2.5-7B-Math2025.12 | 72.6 | |
| Reinforce++Backbone=Qwen2.5-7B-Math2025.12 | 72.4 | |
| CAPO (Reinforce++)Backbone=Qwen2.5-7B-Math2025.12 | 72 | |
| CAPO (GRPO)Backbone=Qwen2.5-1.5B-Math2025.12 | 71.8 | |
| CAPO (RLOO)Backbone=Qwen2.5-1.5B-Math2025.12 | 71.6 | |
| GRPOBackbone=Qwen2.5-1.5B-Math2025.12 | 71.2 | |
| PPOBackbone=Qwen2.5-7B-Math2025.12 | 71 | |
| CAPO (Reinforce++)Backbone=Qwen2.5-1.5B-Math2025.12 | 70.8 | |
| SeqKDTraining Budget=sufficient, Zero-shot=true, Training FLOP=6.6 × 1019, Training Hours=18.32026.02 | 70.8 | |
| CAPO (PPO)Backbone=Qwen2.5-1.5B-Math2025.12 | 70.2 | |
| Reinforce++Backbone=Qwen2.5-1.5B-Math2025.12 | 70 | |
| OPD (prefix scheduling)Training Budget=sufficient, Zero-shot=true, Training FLOP=2.4 × 1019, Training Hours=30.82026.02 | 68.1 | |
| RLOOBackbone=Qwen2.5-1.5B-Math2025.12 | 68 | |
| Qwen2.5-14BBackbone=Qwen2.5-14B, Inference Template=chat inference template2025.05 | 67.4 | |
| OPDTraining Budget=sufficient, Rollout=full, Zero-shot=true, Training FLOP=5.7 × 1019, Training Hours=68.62026.02 | 67.3 | |
| PPOBackbone=Qwen2.5-1.5B-Math2025.12 | 66.6 | |
| Eurus-2-7B-SFT2026.03 | 66.2 | |
| OpenPangu-Embedded-1B + Re2Parameters=1B2026.03 | 65.8 | |
| Llama-3.1-70B-Inst.2026.03 | 65 | |
| M_RLP + PostPre-training Condition=Reinforcement Learning from Pretraining, Post-training=SFT + RLVR2025.09 | 64.3 | |
| Qwen2.5-7BBackbone=Qwen2.5-7B, Inference Template=chat inference template2025.05 | 63.6 | |
| M_CPT + PostPre-training Condition=Continuous Pretraining, Post-training=SFT + RLVR2025.09 | 62.7 | |
| M_base + PostPre-training Condition=Base, Post-training=SFT + RLVR2025.09 | 61.92 | |
| SeqKDSteps=300, Training Budget=limited, Zero-shot=true, Training FLOP=1.3 × 1019, Training Hours=3.72026.02 | 60.5 | |
| CoTBackbone=Qwen2.5-1.5B-Math2025.12 | 59 | |
| M_RLPPre-training Condition=Reinforcement Learning from Pretraining, Post-training=None2025.09 | 58.48 | |
| OPDPrefix Length=2048, Training Budget=limited, Zero-shot=true, Training FLOP=4.7 × 1018, Training Hours=5.02026.02 | 58.3 | |
| M_CPTPre-training Condition=Continuous Pretraining, Post-training=None2025.09 | 57.52 | |
| OPDPrefix Length=1024, Training Budget=limited, Zero-shot=true, Training FLOP=2.5 × 1018, Training Hours=3.42026.02 | 54.5 | |
| CoTBackbone=Qwen2.5-7B-Math2025.12 | 50.8 | |
| OPDPrefix Length=256, Training Budget=limited, Zero-shot=true, Training FLOP=9.4 × 1017, Training Hours=2.42026.02 | 50 | |
| OpenPangu-Embedded-1BParameters=1B2026.03 | 49.3 | |
| M_basePre-training Condition=Base, Post-training=None2025.09 | 48.45 | |
| OPDPrefix Length=512, Training Budget=limited, Zero-shot=true, Training FLOP=1.4 × 1018, Training Hours=2.82026.02 | 47.6 | |
| OPDSteps=10, Training Budget=limited, Zero-shot=true, Training FLOP=8.2 × 1018, Training Hours=12.22026.02 | 43.4 | |
| Qwen3-1.7B-BasePost Training=No, Zero-shot=true2026.02 | 20.1 | |
| MoSLoRABase Model=Qwen2-1.5B, activated parameters=1.31%2026.02 | 18 | |
| HydraLoRABase Model=Qwen2-1.5B, activated parameters=2.54%2026.02 | 16.8 | |
| AdaLoRABase Model=Qwen2-1.5B, rank=82026.02 | 16.6 | |
| MELoRABase Model=Qwen2-1.5B, rank=82026.02 | 16.4 | |
| LoRABase Model=Qwen2-1.5B, rank=82026.02 | 15.6 | |
| HypLoRABase Model=Qwen2-1.5B, rank=82026.02 | 14.6 | |
| DoRABase Model=Qwen2-1.5B, rank=82026.02 | 14.2 | |
| LoRABase Model=Qwen2-1.5B, rank=increased2026.02 | 13.6 | |
| HMoRABase Model=Qwen2-1.5B, activated parameters=2.26%2026.02 | 2.6 | |
| BaseBase Model=Qwen2-1.5B2026.02 | 0 |