General Reasoning on MMLU-Pro
90.1MMLU-Pro General Reasoning Avg@8 AccGemini 3 Pro
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Gemini 3 ProEvaluation Source=External2026.02 | 90.1 | — | — | |
| Claude Opus 4.5Evaluation Source=Internal2026.02 | 89.3 | — | — | |
| Kimi K2.5Evaluation Source=External2026.02 | 87.1 | — | — | |
| GPT-5.2 (xhigh)Evaluation Source=Internal2026.02 | 86.7 | — | — | |
| DeepSeek-V3.2Evaluation Source=External2026.02 | 85 | — | — | |
| SUPERNOVA-4BModel Size=4B2026.04 | 76 | — | 80.1 | |
| rSIM (Plug-in)Reasoner=Llama3.3-70B, Planner=3B plug-in2025.12 | 72.7 | 5 | — | |
| rSIM (Transfer)Reasoner=Llama3.3-70B, Planner=Qwen2.5-1.5B2025.12 | 72.3 | 6 | — | |
| rSIM (Plug-in)Reasoner=Llama3.3-70B, Planner=1B plug-in2025.12 | 71.8 | 6 | — | |
| PromptingReasoner=Llama3.3-70B, Planner=70B2025.12 | 71.5 | 4 | — | |
| Qwen3-4BModel Size=4B2026.04 | 71.2 | — | 71.3 | |
| PS+Reasoner=Llama3.3-70B, Planner=No2025.12 | 70 | 0 | — | |
| ZeroCoTReasoner=Llama3.3-70B, Planner=No2025.12 | 68.9 | 0 | — | |
| PromptingReasoner=Llama3.3-70B, Planner=3B2025.12 | 68.9 | 7 | — | |
| General-ReasonerBackbone=Qwen3-8B-Base, Methodology=General-Reasoner, Supervision Ratio=232k WebInstruct2025.12 | 65.1 | — | — | |
| SPICEBackbone=Qwen3-8B-Base, Methodology=SPICE, Supervision Ratio=None2025.12 | 65 | — | — | |
| Qwen3-1.7BModel Size=1.7B2026.04 | 64.3 | — | 67.8 | |
| R-Few (5%)Backbone=Qwen3-8B-Base, Methodology=R-Few, Supervision Ratio=5%2025.12 | 63.2 | — | — | |
| GRPO-DBBBackbone=Qwen3-8B-Base, Training Dataset=DAPO-Math-17k2026.03 | 63.12 | — | — | |
| General-ReasonerBackbone=Qwen3-4B-Base, Methodology=General-Reasoner, Supervision Ratio=232k WebInstruct2025.12 | 62.8 | — | — | |
| R-Few (1%)Backbone=Qwen3-8B-Base, Methodology=R-Few, Supervision Ratio=1%2025.12 | 62.8 | — | — | |
| Absolute ZeroBackbone=Qwen3-8B-Base, Methodology=Absolute Zero, Supervision Ratio=None2025.12 | 62.5 | — | — | |
| R-ZeroBackbone=Qwen3-8B-Base, Methodology=R-Zero, Supervision Ratio=None2025.12 | 61.6 | — | — | |
| SUPERNOVA-1.7BModel Size=1.7B2026.04 | 61.5 | — | 75.2 | |
| SPICEBackbone=Qwen3-4B-Base, Methodology=SPICE, Supervision Ratio=None2025.12 | 58.1 | — | — | |
| Base ModelBackbone=Qwen3-8B-Base, Methodology=Base, Supervision Ratio=None2025.12 | 58 | — | — | |
| R-Few (5%)Backbone=Qwen3-4B-Base, Methodology=R-Few, Supervision Ratio=5%2025.12 | 56.2 | — | — | |
| SUPERNOVA-0.6BModel Size=0.6B2026.04 | 56.2 | — | 64.6 | |
| R-Few (1%)Backbone=Qwen3-4B-Base, Methodology=R-Few, Supervision Ratio=1%2025.12 | 55.9 | — | — | |
| Qwen3-0.6BModel Size=0.6B2026.04 | 55.3 | — | 53.5 | |
| R-ZeroBackbone=Qwen3-4B-Base, Methodology=R-Zero, Supervision Ratio=None2025.12 | 54.2 | — | — | |
| RePOBackbone=Qwen3-8B-Base, Training Dataset=DAPO-Math-17k2026.03 | 53.15 | — | — | |
| Absolute ZeroBackbone=Qwen3-4B-Base, Methodology=Absolute Zero, Supervision Ratio=None2025.12 | 52.6 | — | — | |
| Base ModelBackbone=Qwen3-4B-Base, Methodology=Base, Supervision Ratio=None2025.12 | 51.6 | — | — | |
| GRPOBackbone=Qwen3-8B-Base, Training Dataset=DAPO-Math-17k2026.03 | 50.02 | — | — | |
| GRPO-DBBBackbone=Qwen3-1.7B-Base, Training Dataset=DAPO-Math-17k2026.03 | 40.83 | — | — | |
| RePOBackbone=Qwen3-1.7B-Base, Training Dataset=DAPO-Math-17k2026.03 | 35.49 | — | — | |
| GRPOBackbone=Qwen3-1.7B-Base, Training Dataset=DAPO-Math-17k2026.03 | 33.72 | — | — | |
| rSIMReasoner=Llama3.2-1B, Planner=3B2025.12 | 33 | 6 | — | |
| PromptingReasoner=Llama3.2-3B, Planner=70B2025.12 | 31.8 | 5 | — | |
| rSIMReasoner=Llama3.2-1B, Planner=Qwen2.5-1.5B2025.12 | 31.8 | 4 | — | |
| rSIMReasoner=Llama3.2-1B, Planner=1B2025.12 | 30.8 | 4 | — | |
| ZeroCoTReasoner=Llama3.2-3B, Planner=No2025.12 | 30.1 | 0 | — | |
| PS+Reasoner=Llama3.2-3B, Planner=No2025.12 | 30 | 0 | — | |
| PromptingReasoner=Llama3.2-3B, Planner=3B2025.12 | 28.5 | 7 | — | |
| PromptingReasoner=Llama3.2-1B, Planner=70B2025.12 | 22 | 6 | — | |
| ZeroCoTReasoner=Llama3.2-1B, Planner=No2025.12 | 21.2 | 0 | — | |
| PS+Reasoner=Llama3.2-1B, Planner=No2025.12 | 19.4 | 0 | — | |
| PromptingReasoner=Llama3.2-1B, Planner=3B2025.12 | 16.8 | 8 | — | |
| GPSBackbone=DeepSeek-R1-Distill-7B, Strategy=GPS, Runtime=49h2026.02 | 0.528 | — | — | |
| MoPPSBackbone=DeepSeek-R1-Distill-7B, Strategy=MoPPS, Runtime=42h2026.02 | 0.524 | — | — | |
| GRESOBackbone=DeepSeek-R1-Distill-7B, Strategy=GRESO, Runtime=53h2026.02 | 0.522 | — | — | |
| Dynamic Sampling (Oracle)Backbone=DeepSeek-R1-Distill-7B, Strategy=DS, Runtime=77h2026.02 | 0.52 | — | — | |
| PCLBackbone=DeepSeek-R1-Distill-7B, Strategy=PCL, Runtime=50h2026.02 | 0.506 | — | — | |
| DeepSeek-R1-Distill-7BBackbone=DeepSeek-R1-Distill-7B, Strategy=Base2026.02 | 0.505 | — | — | |
| Uniform SamplingBackbone=DeepSeek-R1-Distill-7B, Strategy=Uniform, Runtime=40h2026.02 | 0.498 | — | — | |
| GRESOBackbone=DeepSeek-R1-Distill-1.5B, Strategy=GRESO, Runtime=27h2026.02 | 0.247 | — | — | |
| Dynamic Sampling (Oracle)Backbone=DeepSeek-R1-Distill-1.5B, Strategy=DS, Runtime=30h2026.02 | 0.245 | — | — | |
| PCLBackbone=DeepSeek-R1-Distill-1.5B, Strategy=PCL, Runtime=17h2026.02 | 0.24 | — | — | |
| GPSBackbone=DeepSeek-R1-Distill-1.5B, Strategy=GPS, Runtime=16h2026.02 | 0.233 | — | — | |
| Uniform SamplingBackbone=DeepSeek-R1-Distill-1.5B, Strategy=Uniform, Runtime=16h2026.02 | 0.223 | — | — | |
| MoPPSBackbone=DeepSeek-R1-Distill-1.5B, Strategy=MoPPS, Runtime=17h2026.02 | 0.214 | — | — | |
| DeepSeek-R1-Distill-1.5BBackbone=DeepSeek-R1-Distill-1.5B, Strategy=Base2026.02 | 0.212 | — | — |