Mathematical Reasoning on Olympiad (test)
52.1AccuracyOpenAI-o1-preview
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| OpenAI-o1-previewCode Integration=No2025.02 | 52.1 | — | |
| GPT-4oCode Integration=No2025.02 | 43.3 | — | |
| NuminaMath-72BCode Integration=Yes2025.02 | 36.7 | — | |
| AutoCode4Math-Qwen2.5Code Integration=Autonomous2025.02 | 32.6 | — | |
| Qwen-2.5-Base-7BCode Integration=No2025.02 | 30.37 | — | |
| AutoCode4Math-DeepSeekCode Integration=Autonomous2025.02 | 26.95 | — | |
| AutoCode4Math-Qwen2Code Integration=Autonomous2025.02 | 26.37 | — | |
| DeepSeekMath-CRPS-60KBackbone=DeepSeekMath-7B, #Samples=60K2026.04 | 24.6 | — | |
| MathFusionBackbone=DeepSeekMath-7B, #Samples=60K2026.04 | 23.3 | — | |
| DeepSeekMath-CRPS-30KBackbone=DeepSeekMath-7B, #Samples=30K2026.04 | 23.2 | — | |
| Dart-Math-Llama3-8BCode Integration=No2025.02 | 23 | — | |
| SIGMA-60KBackbone=DeepSeekMath-7B, #Samples=60K2026.04 | 22.5 | — | |
| NuminaMath-7B-CoTCode Integration=No2025.02 | 22.22 | — | |
| DeepSeekMath-DARTBackbone=DeepSeekMath-7B, #Samples=590K2026.04 | 21.7 | — | |
| Qwen2Math-Base-7BCode Integration=No2025.02 | 21.62 | — | |
| MathFusion (Sequential)Backbone=DeepSeekMath-7B, #Samples=30K2026.04 | 21.6 | — | |
| SIGMA-30KBackbone=DeepSeekMath-7B, #Samples=30K2026.04 | 21.6 | — | |
| Mathstral-7BCode Integration=No2025.02 | 21.5 | — | |
| DeepSeekMath-CRPS-15KBackbone=DeepSeekMath-7B, #Samples=15K2026.04 | 21.4 | — | |
| DeepSeekMath-DARTBackbone=DeepSeekMath-7B, #Samples=60K2026.04 | 21 | — | |
| DeepseekMath-Instruct-7BCode Integration=Yes2025.02 | 20.44 | — | |
| DeepSeekMath-RFTBackbone=DeepSeekMath-7B, #Samples=590K2026.04 | 19.1 | — | |
| Dart-Math-DeepSeek-7BCode Integration=No2025.02 | 18.52 | — | |
| DeepSeekMath-InstructBackbone=DeepSeekMath-7B, #Samples=780K2026.04 | 14.2 | — | |
| DeepSeekMath-MMIQCBackbone=DeepSeekMath-7B, #Samples=2.3M2026.04 | 13 | — | |
| Mammoth-Mistral-7BCode Integration=Yes2025.02 | 9.63 | — | |
| DeepSeekMath-MetaMathBackbone=DeepSeekMath-7B, #Samples=60K2026.04 | 9.5 | — | |
| DAPOModel Category=RL Post-Trained Models2026.05 | — | 41.8 | |
| GRPOModel Category=RL Post-Trained Models2026.05 | — | 40.5 | |
| GRPOBackbone=Qwen3-4B-Base, Context Length=1K2025.10 | — | 42.8 | |
| GRPOBackbone=Qwen3-8B-Base, Context Length=1K2025.10 | — | 44.2 | |
| GRPOBackbone=Qwen3-4B-Base, Context Length=8K2025.10 | — | 49.9 | |
| GRPO + coupled rhythm creditBackbone=Qwen3-4B-Base, Context Length=1K2025.10 | — | 44.1 | |
| GRPO + coupled rhythm creditBackbone=Qwen3-8B-Base, Context Length=1K2025.10 | — | 47 | |
| GRPO + coupled rhythm creditBackbone=Qwen3-4B-Base, Context Length=8K2025.10 | — | 52.2 | |
| GRPO + global-anchor creditBackbone=Qwen3-4B-Base, Context Length=1K2025.10 | — | 43 | |
| GRPO + global-anchor creditBackbone=Qwen3-8B-Base, Context Length=1K2025.10 | — | 46.1 | |
| GRPO + global-anchor creditBackbone=Qwen3-4B-Base, Context Length=8K2025.10 | — | 51.2 | |
| GRPO + high-entropy creditBackbone=Qwen3-4B-Base, Context Length=1K2025.10 | — | 42.5 | |
| GRPO + high-entropy creditBackbone=Qwen3-8B-Base, Context Length=1K2025.10 | — | 45.6 | |
| GRPO + high-entropy creditBackbone=Qwen3-4B-Base, Context Length=8K2025.10 | — | 48.6 | |
| GRPO + local-chunk creditBackbone=Qwen3-4B-Base, Context Length=1K2025.10 | — | 43.1 | |
| GRPO + local-chunk creditBackbone=Qwen3-8B-Base, Context Length=1K2025.10 | — | 45.9 | |
| GRPO + local-chunk creditBackbone=Qwen3-4B-Base, Context Length=8K2025.10 | — | 51.5 | |
| GRPO + random creditBackbone=Qwen3-4B-Base, Context Length=1K2025.10 | — | 42 | |
| GRPO + random creditBackbone=Qwen3-8B-Base, Context Length=1K2025.10 | — | 43.3 | |
| GRPO + random creditBackbone=Qwen3-4B-Base, Context Length=8K2025.10 | — | 50 | |
| GSPOModel Category=RL Post-Trained Models2026.05 | — | 39.8 | |
| HölderPOModel Category=RL Post-Trained Models, Schedule=Linear Des: 2 → −22026.05 | — | 40.6 | |
| Qwen3-4B-BaseModel Category=Base Model2026.05 | — | 28.6 |