Scientific Reasoning on MMLU-STEM
87.4AccuracyGPT-4o
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-4on=3,014, Shots (Few-shot)=5s, Iterative Repair Setting=no iterative repair2026.04 | 87.4 | — | |
| AeroTherm-GPTn=3,014, Shots (Few-shot)=5s, Iterative Repair Setting=no iterative repair2026.04 | 80.7 | — | |
| Llama-3-70Bn=3,014, Shots (Few-shot)=5s, Iterative Repair Setting=no iterative repair2026.04 | 79.2 | — | |
| SuCoModel Scale=7B2026.06 | 75.8 | — | |
| DeepCompress-Zero-7BParameters=7B2025.10 | 75.5 | — | |
| DeepMath-Zero-7BParameters=7B2025.10 | 72.7 | — | |
| Qwen2.5-7BModel Size=>= 7B, Reasoning Protocol=few-shot CoT2025.12 | 72.3 | 63.5 | |
| S-GRPOModel Scale=7B2026.06 | 72.2 | — | |
| LHRMsModel Scale=7B2026.06 | 71.4 | — | |
| AdaCoTModel Scale=7B2026.06 | 69 | — | |
| AdaptThinkModel Scale=7B2026.06 | 68.8 | — | |
| DeepSeek-R1-DistillModel Scale=7B2026.06 | 67.6 | — | |
| Gemma-2-9BModel Size=>= 7B, Reasoning Protocol=few-shot CoT2025.12 | 65.1 | 55.6 | |
| MAmmoTH2-8BModel Size=>= 7B, Reasoning Protocol=few-shot CoT2025.12 | 64.2 | 65.2 | |
| MAmmoTH2-7BModel Size=>= 7B, Reasoning Protocol=few-shot CoT2025.12 | 62.4 | 57.8 | |
| M3POBackbone=Qwen2.5-3B-Instruct2025.12 | 61.6 | 70.5 | |
| GRPOBackbone=Qwen2.5-3B-Instruct2025.12 | 60.1 | 69.1 | |
| HRPOBackbone=Qwen2.5-3B-Instruct2025.12 | 59 | 68.7 | |
| PPOBackbone=Qwen2.5-3B-Instruct2025.12 | 58.2 | 68.2 | |
| M3POBackbone=Qwen2.5-1.5B-Instruct2025.12 | 58.1 | 60.3 | |
| HRPOBackbone=Qwen2.5-1.5B-Instruct2025.12 | 56.9 | 58.1 | |
| PPOBackbone=Qwen2.5-1.5B-Instruct2025.12 | 56.6 | 57.2 | |
| DeepSeekMath-7BModel Size=>= 7B, Reasoning Protocol=few-shot CoT2025.12 | 56.5 | 51.9 | |
| GRPOBackbone=Qwen2.5-1.5B-Instruct2025.12 | 56.2 | 57.9 | |
| DeepCompress-Zero-3BParameters=3B2025.10 | 56 | — | |
| DeepMath-Zero-3BParameters=3B2025.10 | 54.7 | — | |
| Math-InstructModel Scale=7B2026.06 | 51.3 | — | |
| Qwen-2.5-3B-InstructParameters=3B, Variant=Instruct2025.10 | 48.7 | — | |
| Open-Reasoner-Zero-7BParameters=7B2025.10 | 47 | — | |
| SFTBackbone=Qwen2.5-3B-Instruct2025.12 | 45.4 | 46.1 | |
| SFTBackbone=Qwen2.5-1.5B-Instruct2025.12 | 40.3 | 43.3 | |
| SuCoModel Scale=1.5B2026.06 | 38.6 | — | |
| LHRMsModel Scale=1.5B2026.06 | 35.7 | — | |
| S-GRPOModel Scale=1.5B2026.06 | 34.8 | — | |
| AdaptThinkModel Scale=1.5B2026.06 | 34.4 | — | |
| AdaCoTModel Scale=1.5B2026.06 | 33.9 | — | |
| DeepSeek-R1-DistillModel Scale=1.5B2026.06 | 33.8 | — | |
| Math-InstructModel Scale=1.5B2026.06 | 30.2 | — | |
| Math-BaseModel Scale=7B2026.06 | 26.4 | — | |
| Qwen-2.5-7B-SimpleRL-ZooParameters=7B, Variant=SimpleRL-Zoo2025.10 | 15 | — | |
| Math-BaseModel Scale=1.5B2026.06 | 14.5 | — | |
| Qwen-2.5-7BParameters=7B2025.10 | 12.1 | — | |
| Qwen-2.5-3BParameters=3B2025.10 | 7.5 | — |