Mathematical Reasoning on MATH 500 (accuracy)
97.9MATH 500 Accuracyo3-mini-high
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| o3-mini-highaccess=API only, examples_for_finetuning=N/A2025.01 | 97.9 | — | — | — | |
| o3-mini-mediumaccess=API only, examples_for_finetuning=N/A2025.01 | 97.3 | — | — | — | |
| r1access=Open Weights, examples_for_finetuning=>800K2025.01 | 97.3 | — | — | — | |
| Gemini 2.5 proEvaluation Protocol=pass@12025.09 | 96.7 | — | — | — | |
| o3-mini-lowaccess=API only, examples_for_finetuning=N/A2025.01 | 95.8 | — | — | — | |
| s1.1access=Open Weights and Open Data, examples_for_finetuning=1K, budget_forcing=Wait 1x2025.01 | 95.4 | — | — | — | |
| s1.1access=Open Weights and Open Data, examples_for_finetuning=1K, budget_forcing=Wait 2x2025.01 | 95.4 | — | — | — | |
| PACS-8BEvaluation Protocol=pass@1, Backbone=Qwen3-8B2025.09 | 95.09 | — | — | — | |
| LIMOaccess=Open Weights and Open Data, examples_for_finetuning=8172025.01 | 94.8 | — | — | — | |
| PACS-4BEvaluation Protocol=pass@1, Backbone=Qwen3-4B2025.09 | 94.8 | — | — | — | |
| r1-distill-Llama-70Baccess=Open Weights, examples_for_finetuning=800K2025.01 | 94.5 | — | — | — | |
| s1.1access=Open Weights and Open Data, examples_for_finetuning=1K, budget_forcing=None2025.01 | 94.4 | — | — | — | |
| r1-distill-Qwen-32Baccess=Open Weights, examples_for_finetuning=800K2025.01 | 94.3 | — | — | — | |
| r1-distill-Qwen-14Baccess=Open Weights, examples_for_finetuning=800K2025.01 | 93.9 | — | — | — | |
| s1access=Open Weights and Open Data, examples_for_finetuning=1K, budget_forcing=Wait 2x2025.01 | 93 | — | — | — | |
| s1access=Open Weights and Open Data, examples_for_finetuning=1K, budget_forcing=Wait 1x2025.01 | 92.8 | — | — | — | |
| s1access=Open Weights and Open Data, examples_for_finetuning=1K, budget_forcing=None2025.01 | 92.6 | — | — | — | |
| s1access=Open Weights and Open Data, examples_for_finetuning=1K, budget_forcing=Wait 4x2025.01 | 92.2 | — | — | — | |
| Qwen3-14BEvaluation Protocol=pass@12025.09 | 92.07 | — | — | — | |
| Qwen3-4BEvaluation Protocol=pass@12025.09 | 91.01 | — | — | — | |
| DS-R1-Distill-Qwen-7B-ScaleQuestModel=DS-R1-Distill-Qwen-7B-ScaleQuest, Data=ScaleQuest2024.10 | 91 | — | — | — | |
| LIMOEvaluation Protocol=pass@12025.09 | 91 | — | — | — | |
| QwQ-32Baccess=Open Weights, examples_for_finetuning=N.A.2025.01 | 90.6 | — | — | — | |
| DS-R1-Distill-Qwen-7BModel=DS-R1-Distill-Qwen-7B2024.10 | 90.2 | — | — | — | |
| ESPO2025.11 | 90.2 | — | — | — | |
| Qwen3-8BEvaluation Protocol=pass@12025.09 | 89.94 | — | — | — | |
| ARLCPBase Model=Qwen3-1.7B2026.02 | 89.6 | 2,568 | — | — | |
| ARLCPBase Model=DeepSeek-R1-Distill-Llama-8B2026.02 | 89 | 2,080 | — | — | |
| OpenThinker3-7BEvaluation Protocol=pass@12025.09 | 88.9 | — | — | — | |
| VanillaBase Model=DeepSeek-R1-Distill-Llama-8B2026.02 | 87.6 | 4,062 | — | — | |
| VanillaBase Model=Qwen3-1.7B2026.02 | 87 | 5,096 | — | — | |
| Ouro2.6B-Thinking + RLTTDecoding Method=Deterministic, Evaluation Protocol=Zero-shot, Training=RLTT2026.02 | 86 | — | — | — | |
| DAPO2025.11 | 85.8 | — | — | — | |
| Qwen3-4BTotal Parameters=4B, Active Parameters=4B, Trained Tokens=36T2025.11 | 85.6 | — | — | — | |
| GSPO2025.11 | 85.4 | — | — | — | |
| Sky-T1-7BEvaluation Protocol=pass@12025.09 | 85.38 | — | — | — | |
| PRIME-Zero-7BParameter size=7B2026.02 | 83.8 | — | — | — | |
| Eurus-7BParameter size=7B2026.02 | 83.8 | — | — | — | |
| GRPOBackbone=Qwen3-4B-Base, Sample Count=Avg@4, Evaluation Domain=In-domain2025.12 | 83.6 | — | — | — | |
| ReProBackbone=Qwen3-4B-Base, Sample Count=Avg@4, Evaluation Domain=In-domain, Base Method=GRPO2025.12 | 83 | — | — | — | |
| SAGEBase Model=Qwen2.5-7B-Instruct2026.02 | 82.8 | — | — | — | |
| COMPACTStudent Architecture=Qwen2.5-7B, Distillation Method=COMPACT, Evaluation Protocol=Fine-tuned2026.01 | 82.7 | — | — | — | |
| OpenReasoner-Zero-7B @ 8kParameter size=7B, Inference budget=8k2026.02 | 82.4 | — | — | — | |
| MCC-KDStudent Architecture=Qwen2.5-7B, Distillation Method=MCC-KD, Evaluation Protocol=Fine-tuned2026.01 | 82.2 | — | — | — | |
| Eurus-2-7B-PRIMEEvaluation Protocol=pass@12025.09 | 82.07 | — | — | — | |
| DPO (Full)Base Model=Qwen2.5-7B-Instruct2026.02 | 82 | — | — | — | |
| VanillaBase Model=Qwen2.5-7B-Instruct2026.02 | 81.6 | — | — | — | |
| MoTStudent Architecture=Qwen2.5-7B, Distillation Method=MoT, Evaluation Protocol=Fine-tuned2026.01 | 81.6 | — | — | — | |
| DEEP-GRPO-7BParameter size=7B2026.02 | 81.6 | — | — | — | |
| GPG-7BParameter size=7B2026.02 | 80 | — | — | — | |
| Oat-Zero-7B (Dr. GRPO)Parameter size=7B2026.02 | 80 | — | — | — | |
| EDITStudent Architecture=Qwen2.5-7B, Distillation Method=EDIT, Evaluation Protocol=Fine-tuned2026.01 | 79.5 | — | — | — | |
| DPO (Random)Base Model=Qwen2.5-7B-Instruct2026.02 | 79.4 | — | — | — | |
| OpenReasoner-Zero-7B @ 3kParameter size=7B, Inference budget=3k2026.02 | 79.2 | — | — | — | |
| SimpleRL-Zero-7BParameter size=7B2026.02 | 78.2 | — | — | — | |
| SBS-KDStudent Architecture=Qwen2.5-7B, Distillation Method=SBS-KD, Evaluation Protocol=Fine-tuned2026.01 | 77.4 | — | — | — | |
| CommitteeStudent Architecture=Qwen2.5-7B, Distillation Method=Committee, Evaluation Protocol=Fine-tuned2026.01 | 77.25 | — | — | — | |
| Qwen2.5-7B-InstructStudent Architecture=Qwen2.5-7B, Distillation Method=Zero-shot, Evaluation Protocol=Zero-shot2026.01 | 77.2 | — | — | — | |
| DEEP-GRPO-1.5BParameter size=1.5B2026.02 | 75.2 | — | — | — | |
| LFM2-8B-A1BTotal Parameters=8.3B, Active Parameters=1.5B, Trained Tokens=13T2025.11 | 74.2 | — | — | — | |
| Qwen2.5-Math-1.5B-InstructParameter size=1.5B2026.02 | 74.2 | — | — | — | |
| Oat-Zero-1.5B (Dr. GRPO)Parameter size=1.5B2026.02 | 74.2 | — | — | — | |
| SmolLM3-3BTotal Parameters=3.1B, Active Parameters=3.1B, Trained Tokens=11T2025.11 | 73.6 | — | — | — | |
| Gemma-3-4BTotal Parameters=4B, Active Parameters=4B, Trained Tokens=4T2025.11 | 73.2 | — | — | — | |
| Ouro2.6B-Thinking + GRPODecoding Method=Deterministic, Evaluation Protocol=Zero-shot, Training=GRPO2026.02 | 71.6 | — | — | — | |
| Marco-o1Evaluation Protocol=pass@12025.09 | 70.48 | — | — | — | |
| OriginalBackbone=Qwen3-4B-Base, Sample Count=Avg@4, Evaluation Domain=In-domain2025.12 | 68.3 | — | — | — | |
| Ouro2.6B-ThinkingDecoding Method=Deterministic, Evaluation Protocol=Zero-shot2026.02 | 67.8 | — | — | — | |
| SAGEBase Model=Qwen2.5-3B-Instruct2026.02 | 66 | — | — | — | |
| DPO (Full)Base Model=Qwen2.5-3B-Instruct2026.02 | 65.6 | — | — | — | |
| VanillaBase Model=Qwen2.5-3B-Instruct2026.02 | 65.2 | — | — | — | |
| DPO (Random)Base Model=Qwen2.5-3B-Instruct2026.02 | 65.2 | — | — | — | |
| LFM2-2.6BTotal Parameters=2.6B, Active Parameters=2.6B, Trained Tokens=11T2025.11 | 63.6 | — | — | — | |
| Qwen2.5-Math-1.5BParameter size=1.5B2026.02 | 61.8 | — | — | — | |
| DeepSeekR1 – 7BDecoding Method=Deterministic, Evaluation Protocol=Zero-shot2026.02 | 60 | — | — | — | |
| Ouro2.6B-Thinking + SFTDecoding Method=Deterministic, Evaluation Protocol=Zero-shot, Training=SFT2026.02 | 58.2 | — | — | — | |
| Granite-4.0-HTotal Parameters=7B, Active Parameters=1B, Trained Tokens=15T2025.11 | 58.2 | — | — | — | |
| CommitteeStudent Architecture=Qwen2.5-1.5B, Distillation Method=Committee, Evaluation Protocol=Fine-tuned2026.01 | 58.1 | — | — | — | |
| SAGEBase Model=Qwen2.5-1.5B-Instruct2026.02 | 57.2 | — | — | — | |
| DPO (Random)Base Model=Qwen2.5-1.5B-Instruct2026.02 | 56.4 | — | — | — | |
| DPO (Full)Base Model=Qwen2.5-1.5B-Instruct2026.02 | 56.2 | — | — | — | |
| COMPACTStudent Architecture=Qwen2.5-1.5B, Distillation Method=COMPACT, Evaluation Protocol=Fine-tuned2026.01 | 56.2 | — | — | — | |
| DeepSeekR1 – 1.5BDecoding Method=Deterministic, Evaluation Protocol=Zero-shot2026.02 | 55.8 | — | — | — | |
| VanillaBase Model=Qwen2.5-1.5B-Instruct2026.02 | 54.6 | — | — | — | |
| MoTStudent Architecture=Qwen2.5-1.5B, Distillation Method=MoT, Evaluation Protocol=Fine-tuned2026.01 | 54 | — | — | — | |
| MCC-KDStudent Architecture=Qwen2.5-1.5B, Distillation Method=MCC-KD, Evaluation Protocol=Fine-tuned2026.01 | 53.6 | — | — | — | |
| COMPACTStudent Architecture=Llama3.1-8B, Distillation Method=COMPACT, Evaluation Protocol=Fine-tuned2026.01 | 53.56 | — | — | — | |
| MCC-KDStudent Architecture=Llama3.1-8B, Distillation Method=MCC-KD, Evaluation Protocol=Fine-tuned2026.01 | 53 | — | — | — | |
| EDITStudent Architecture=Llama3.1-8B, Distillation Method=EDIT, Evaluation Protocol=Fine-tuned2026.01 | 52.8 | — | — | — | |
| SBS-KDStudent Architecture=Llama3.1-8B, Distillation Method=SBS-KD, Evaluation Protocol=Fine-tuned2026.01 | 52.6 | — | — | — | |
| CommitteeStudent Architecture=Llama3.1-8B, Distillation Method=Committee, Evaluation Protocol=Fine-tuned2026.01 | 51.93 | — | — | — | |
| EDITStudent Architecture=Qwen2.5-1.5B, Distillation Method=EDIT, Evaluation Protocol=Fine-tuned2026.01 | 51.2 | — | — | — | |
| MoTStudent Architecture=Llama3.1-8B, Distillation Method=MoT, Evaluation Protocol=Fine-tuned2026.01 | 50.8 | — | — | — | |
| Qwen2.5-Math-7BParameter size=7B2026.02 | 50.6 | — | — | — | |
| Qwen2.5-1.5B-InstructStudent Architecture=Qwen2.5-1.5B, Distillation Method=Zero-shot, Evaluation Protocol=Zero-shot2026.01 | 50.4 | — | — | — | |
| SBS-KDStudent Architecture=Qwen2.5-1.5B, Distillation Method=SBS-KD, Evaluation Protocol=Fine-tuned2026.01 | 49.8 | — | — | — | |
| Qwen3 – 4BDecoding Method=Deterministic, Evaluation Protocol=Zero-shot2026.02 | 48.4 | — | — | — | |
| Llama3.1-8B-InstructStudent Architecture=Llama3.1-8B, Distillation Method=Zero-shot, Evaluation Protocol=Zero-shot2026.01 | 46 | — | — | — | |
| Deepthought-8BEvaluation Protocol=pass@12025.09 | 45.07 | — | — | — | |
| Llama-3.2-3BTotal Parameters=3.2B, Active Parameters=3.2B, Trained Tokens=9T2025.11 | 41.2 | — | — | — |