Scientific Reasoning on MMLU-Pro (pass@1, avg@10)
80.3Pass@1o1-mini
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| o1-miniModel Category=Thinking Models, Thinking Mode=true2026.02 | 80.3 | — | |
| Dr.SCI-4B-thinkModel Category=Thinking Models, Thinking Mode=true, Model Scale=4B2026.02 | 75.6 | — | |
| GPT-4oModel Category=Instruct Models2026.02 | 74.6 | — | |
| Qwen3-14B-MegaScienceModel Category=Instruct Models, Model Scale=14B, Method Variant=MegaScience2026.02 | 71.9 | — | |
| R1-0528-Qwen3-8BModel Category=Thinking Models, Thinking Mode=true, Model Scale=8B2026.02 | 71.4 | — | |
| Dr.SCI-4B-instructModel Category=Instruct Models, Model Scale=4B2026.02 | 71 | — | |
| Qwen3-4B thinkingModel Category=Thinking Models, Thinking Mode=true, Model Scale=4B2026.02 | 70.4 | — | |
| General-Reasoner-Qw3-14BModel Category=Instruct Models, Model Scale=14B2026.02 | 70.3 | — | |
| R1-Distill-Qwen-32BModel Category=Thinking Models, Thinking Mode=true, Model Scale=32B2026.02 | 67.5 | — | |
| Qwen3-8B-MegaScienceModel Category=Instruct Models, Model Scale=8B, Method Variant=MegaScience2026.02 | 67.3 | — | |
| Qwen3-8B-VeriFreeModel Category=Instruct Models, Model Scale=8B, Method Variant=VeriFree2026.02 | 67.2 | — | |
| QwQ-32BModel Category=Thinking Models, Thinking Mode=true, Model Scale=32B2026.02 | 66.2 | — | |
| Qwen3-4B-VeriFreeModel Category=Instruct Models, Model Scale=4B, Method Variant=VeriFree2026.02 | 63.5 | — | |
| General-Reasoner-4BModel Category=Instruct Models, Model Scale=4B2026.02 | 62.8 | — | |
| Qwen3-4B-MegaScienceModel Category=Instruct Models, Model Scale=4B, Method Variant=MegaScience2026.02 | 61.2 | — | |
| Qwen3-4B non-thinkingModel Category=Instruct Models, Thinking Mode=false, Model Scale=4B2026.02 | 58 | — | |
| Outcome RewardBackbone=Qwen3-4B2025.10 | 57.1 | — | |
| Ours (Sparse)Backbone=Qwen3-4B2025.10 | 55.6 | — | |
| Ours (Dense)Backbone=Qwen3-4B2025.10 | 55.1 | — | |
| SFTBackbone=Qwen3-4B2025.10 | 53.9 | — | |
| Outcome RewardBackbone=Qwen2.5-7B2025.10 | 53.5 | — | |
| Ours (Interval)Backbone=Qwen3-4B2025.10 | 53.5 | — | |
| Qwen3-4B-BaseModel Category=Base, Model Scale=4B2026.02 | 50.6 | — | |
| Ours (Interval)Backbone=Qwen2.5-7B2025.10 | 50.6 | — | |
| Ours (Sparse)Backbone=Qwen2.5-7B2025.10 | 48.5 | — | |
| Outcome RewardBackbone=Llama3.1-8B2025.10 | 48.4 | — | |
| SFTBackbone=Qwen2.5-7B2025.10 | 48.1 | — | |
| SFTBackbone=Llama3.1-8B2025.10 | 47.2 | — | |
| Ours (Dense)Backbone=Qwen2.5-7B2025.10 | 43.8 | — | |
| Ours (Sparse)Backbone=Llama3.1-8B2025.10 | 43.3 | — | |
| Ours (Dense)Backbone=Llama3.1-8B2025.10 | 37.9 | — | |
| Ours (Interval)Backbone=Llama3.1-8B2025.10 | 36.6 | — |