Scientific Reasoning on SuperGPQA
50.1Mean@1Agentic Proposing
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Agentic ProposingTraining Data=Agentic-Proposer-4B Generated Problems, Training Budget=10,000 trajectories, Optimizer=GRPO2026.02 | 50.1 | — | |
| Dr.SCI-4B-thinkModel Category=Thinking Models, Thinking Mode=true, Model Scale=4B2026.02 | 45.7 | — | |
| o1-miniModel Category=Thinking Models, Thinking Mode=true2026.02 | 45.2 | — | |
| Kimi-K2Model Variant=Base, # Shots=5-shot, # Activated Params=32B, # Total Params=1043B2026.02 | 44.7 | — | |
| GPT-4oModel Category=Instruct Models2026.02 | 44.4 | — | |
| Qwen3-14B-MegaScienceModel Category=Instruct Models, Model Scale=14B, Method Variant=MegaScience2026.02 | 44.4 | — | |
| DeepSeek V3.2Model Variant=Exp Base, # Shots=5-shot, # Activated Params=37B, # Total Params=671B2026.02 | 43.6 | — | |
| QwQ-32BModel Category=Thinking Models, Thinking Mode=true, Model Scale=32B2026.02 | 43.6 | — | |
| Dr.SCI-4B-instructModel Category=Instruct Models, Model Scale=4B2026.02 | 43.6 | — | |
| Qwen3-4B-Instruct-2507Protocol=zero-shot2026.02 | 42.8 | — | |
| Qwen3-4B thinkingModel Category=Thinking Models, Thinking Mode=true, Model Scale=4B2026.02 | 42.7 | — | |
| DeepSeek V3.1Model Variant=Base, # Shots=5-shot, # Activated Params=37B, # Total Params=671B2026.02 | 42.3 | — | |
| R1-0528-Qwen3-8BModel Category=Thinking Models, Thinking Mode=true, Model Scale=8B2026.02 | 42.1 | — | |
| MiMo-V2 FlashModel Variant=Base, # Shots=5-shot, # Activated Params=15B, # Total Params=309B2026.02 | 41.1 | — | |
| Step 3.5 FlashModel Variant=Base, # Shots=5-shot, # Activated Params=11B, # Total Params=196B2026.02 | 41 | — | |
| General-Reasoner-Qw3-14BModel Category=Instruct Models, Model Scale=14B2026.02 | 39.9 | — | |
| R1-Distill-Qwen-32BModel Category=Thinking Models, Thinking Mode=true, Model Scale=32B2026.02 | 39.3 | — | |
| Qwen3-8B-MegaScienceModel Category=Instruct Models, Model Scale=8B, Method Variant=MegaScience2026.02 | 38.8 | — | |
| Qwen3-8B-VeriFreeModel Category=Instruct Models, Model Scale=8B, Method Variant=VeriFree2026.02 | 38 | — | |
| Qwen3-4B-VeriFreeModel Category=Instruct Models, Model Scale=4B, Method Variant=VeriFree2026.02 | 35.1 | — | |
| Qwen3-4B-MegaScienceModel Category=Instruct Models, Model Scale=4B, Method Variant=MegaScience2026.02 | 33.1 | — | |
| General-Reasoner-4BModel Category=Instruct Models, Model Scale=4B2026.02 | 32.5 | — | |
| Qwen3-4B non-thinkingModel Category=Instruct Models, Thinking Mode=false, Model Scale=4B2026.02 | 32 | — | |
| SFTBackbone=Qwen2.5-7B-Instruct2025.08 | 29.02 | — | |
| Qwen3-4B-BaseModel Category=Base, Model Scale=4B2026.02 | 28.5 | — | |
| PSFTwarm-upBackbone=Qwen2.5-7B-Instruct2025.08 | 28.37 | — | |
| PSFTBackbone=Qwen2.5-7B-Instruct2025.08 | 28.1 | — | |
| SFT-KLBackbone=Qwen2.5-7B-Instruct2025.08 | 27.69 | — | |
| BaseBackbone=Qwen2.5-7B-Instruct2025.08 | 27.59 | — | |
| PSFTBackbone=Llama3.1-8B-Instruct2025.08 | 25.85 | — | |
| PSFTwarm-upBackbone=Llama3.1-8B-Instruct2025.08 | 25.26 | — | |
| SFTBackbone=Llama3.1-8B-Instruct2025.08 | 21.95 | — | |
| SFT-KLBackbone=Llama3.1-8B-Instruct2025.08 | 19.81 | — | |
| BaseBackbone=Llama3.1-8B-Instruct2025.08 | 18.16 | — |