Scientific Reasoning on Aggregate GPQA, HLE, MMLU-Pro
44.6Average ScoreDr.SCI-4B-think
Evaluation Results
| Method | Links | |
|---|---|---|
| Dr.SCI-4B-thinkModel Category=Thinking Models, Thinking Mode=true, Model Scale=4B2026.02 | 44.6 | |
| o1-miniModel Category=Thinking Models, Thinking Mode=true2026.02 | 43.4 | |
| R1-0528-Qwen3-8BModel Category=Thinking Models, Thinking Mode=true, Model Scale=8B2026.02 | 41.8 | |
| R1-Distill-Qwen-32BModel Category=Thinking Models, Thinking Mode=true, Model Scale=32B2026.02 | 41 | |
| Dr.SCI-4B-instructModel Category=Instruct Models, Model Scale=4B2026.02 | 40.2 | |
| GPT-4oModel Category=Instruct Models2026.02 | 39 | |
| Qwen3-14B-MegaScienceModel Category=Instruct Models, Model Scale=14B, Method Variant=MegaScience2026.02 | 39 | |
| Qwen3-4B thinkingModel Category=Thinking Models, Thinking Mode=true, Model Scale=4B2026.02 | 38.9 | |
| General-Reasoner-Qw3-14BModel Category=Instruct Models, Model Scale=14B2026.02 | 38.9 | |
| QwQ-32BModel Category=Thinking Models, Thinking Mode=true, Model Scale=32B2026.02 | 37.4 | |
| Qwen3-8B-MegaScienceModel Category=Instruct Models, Model Scale=8B, Method Variant=MegaScience2026.02 | 35.5 | |
| Qwen3-8B-VeriFreeModel Category=Instruct Models, Model Scale=8B, Method Variant=VeriFree2026.02 | 34.6 | |
| Qwen3-4B-VeriFreeModel Category=Instruct Models, Model Scale=4B, Method Variant=VeriFree2026.02 | 32.3 | |
| General-Reasoner-4BModel Category=Instruct Models, Model Scale=4B2026.02 | 31.6 | |
| Qwen3-4B-MegaScienceModel Category=Instruct Models, Model Scale=4B, Method Variant=MegaScience2026.02 | 29.3 | |
| Qwen3-4B non-thinkingModel Category=Instruct Models, Thinking Mode=false, Model Scale=4B2026.02 | 29.2 | |
| Qwen3-4B-BaseModel Category=Base, Model Scale=4B2026.02 | 24.5 |