Scientific Reasoning on GPQA Diamond (pass@1 and avg@10)
64.58Pass@1Large Model
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Large ModelModel Pair=Qwen3-32 / 1.7B2026.02 | 64.58 | — | |
| RelayGenModel Pair=Qwen3-32 / 1.7B2026.02 | 63.64 | — | |
| R2RModel Pair=Qwen3-32 / 1.7B2026.02 | 61.62 | — | |
| Large ModelModel Pair=R1-Distill-Qwen-32B / R1-Distill-Qwen-1.5B2026.02 | 60.61 | — | |
| RelayGenModel Pair=R1-Distill-Qwen-32B / R1-Distill-Qwen-1.5B2026.02 | 56.82 | — | |
| R2RModel Pair=R1-Distill-Qwen-32B / R1-Distill-Qwen-1.5B2026.02 | 48.99 | — | |
| Critique-GRPO (Self-Critique)w/ External Supervision=true2025.06 | 47.98 | — | |
| Critique-GRPO (Self-Critique & Self-Evaluation)w/ External Supervision=false2025.06 | 47.47 | — | |
| Speculative ThinkingModel Pair=R1-Distill-Qwen-32B / R1-Distill-Qwen-1.5B2026.02 | 42.68 | — | |
| Speculative ThinkingModel Pair=Qwen3-32 / 1.7B2026.02 | 41.29 | — | |
| R1-GRPOw/ External Supervision=true2025.06 | 40.4 | — | |
| SFTw/ External Supervision=true2025.06 | 38.38 | — | |
| Small ModelModel Pair=Qwen3-32 / 1.7B2026.02 | 37.33 | — | |
| Qwen3-8B (w/ Think)2025.06 | 35.86 | — | |
| Small ModelModel Pair=R1-Distill-Qwen-32B / R1-Distill-Qwen-1.5B2026.02 | 30.3 | — | |
| Dr.SCI-4B-thinkModel Category=Thinking Models, Thinking Mode=true, Model Scale=4B2026.02 | 0.632 | — | |
| R1-Distill-Qwen-32BModel Category=Thinking Models, Thinking Mode=true, Model Scale=32B2026.02 | 0.621 | — | |
| R1-0528-Qwen3-8BModel Category=Thinking Models, Thinking Mode=true, Model Scale=8B2026.02 | 0.611 | — | |
| o1-miniModel Category=Thinking Models, Thinking Mode=true2026.02 | 0.6 | — | |
| Dr.SCI-4B-instructModel Category=Instruct Models, Model Scale=4B2026.02 | 0.566 | — | |
| General-Reasoner-Qw3-14BModel Category=Instruct Models, Model Scale=14B2026.02 | 0.561 | — | |
| Qwen3-4B thinkingModel Category=Thinking Models, Thinking Mode=true, Model Scale=4B2026.02 | 0.559 | — | |
| QwQ-32BModel Category=Thinking Models, Thinking Mode=true, Model Scale=32B2026.02 | 0.553 | — | |
| Qwen3-14B-MegaScienceModel Category=Instruct Models, Model Scale=14B, Method Variant=MegaScience2026.02 | 0.505 | — | |
| GPT-4oModel Category=Instruct Models2026.02 | 0.5 | — | |
| Qwen3-8B-MegaScienceModel Category=Instruct Models, Model Scale=8B, Method Variant=MegaScience2026.02 | 0.465 | — | |
| Qwen3-8B-VeriFreeModel Category=Instruct Models, Model Scale=8B, Method Variant=VeriFree2026.02 | 0.444 | — | |
| General-Reasoner-4BModel Category=Instruct Models, Model Scale=4B2026.02 | 0.429 | — | |
| Qwen3-4B-VeriFreeModel Category=Instruct Models, Model Scale=4B, Method Variant=VeriFree2026.02 | 0.424 | — | |
| Qwen3-4B non-thinkingModel Category=Instruct Models, Thinking Mode=false, Model Scale=4B2026.02 | 0.417 | — | |
| Qwen3-4B-BaseModel Category=Base, Model Scale=4B2026.02 | 0.367 | — | |
| Qwen3-4B-MegaScienceModel Category=Instruct Models, Model Scale=4B, Method Variant=MegaScience2026.02 | 0.349 | — |