Scientific Reasoning on PrincipiaBench
48.74RealMath Scoreo3
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| o3Training Data=-2026.03 | 48.74 | 63.75 | 81.91 | 57.19 | 62.9 | |
| GPT-OSS-120BTraining Data=-2026.03 | 44.05 | 59.89 | 74.47 | 53.92 | 58.08 | |
| GPT-OSS-20BTraining Data=-2026.03 | 42.56 | 51.59 | 72.34 | 50.71 | 54.3 | |
| Online Multitask + aggregationBackbone=Qwen3-4B-Instruct-2507, Aggregation=Enabled2026.03 | 37.12 | 60.1 | 79.23 | 58.25 | 58.68 | |
| ParaGator-4B-Instruct-Principia + aggregationBackbone=Qwen3-4B-Instruct-2507, Aggregation=Enabled2026.03 | 36.84 | 59.41 | 79.19 | 56.85 | 58.07 | |
| GPT-4.1Training Data=-2026.03 | 36.3 | 51.25 | 61.44 | 46.43 | 48.85 | |
| Qwen3-235B (thinking)Training Data=-2026.03 | 36.02 | 58.64 | 73.94 | 53.74 | 55.58 | |
| Dr.GRPO (with KL) + aggregationBackbone=Qwen3-4B-Instruct-2507, Aggregation=Enabled2026.03 | 35.44 | 57.14 | 76.67 | 58.94 | 57.05 | |
| Offline aggregation (AggLM) + aggregationBackbone=Qwen3-4B-Instruct-2507, Aggregation=Enabled2026.03 | 33.19 | 51.04 | 72.13 | 52.3 | 52.17 | |
| Claude-4.0-SonnetTraining Data=-2026.03 | 32.04 | 41.82 | 59.57 | 48.19 | 45.4 | |
| Qwen3-4B-Instruct-2507 + aggregationBackbone=Qwen3-4B-Instruct-2507, Aggregation=Enabled2026.03 | 29.45 | 44.12 | 72.56 | 49.09 | 48.8 | |
| Online MultitaskBackbone=Qwen3-4B-Instruct-2507, Aggregation=None2026.03 | 29.16 | 47.07 | 67.65 | 53.44 | 49.32 | |
| Principia-4BTraining Data=Principia Collec.2026.03 | 28.96 | 51.24 | 66.53 | 47.05 | 48.45 | |
| Qwen3-235B (no-thinking)Training Data=-2026.03 | 28.54 | 44.32 | 63.03 | 45.14 | 45.26 | |
| Qwen3-14B (thinking)Training Data=-2026.03 | 28.36 | 51.36 | 67.02 | 49.35 | 49.02 | |
| Dr.GRPO (with KL)Backbone=Qwen3-4B-Instruct-2507, Aggregation=None2026.03 | 28.13 | 49.95 | 68.74 | 52.31 | 49.78 | |
| ParaGator-4B-Instruct-PrincipiaBackbone=Qwen3-4B-Instruct-2507, Aggregation=None2026.03 | 27.93 | 46.18 | 66.7 | 52.39 | 48.3 | |
| ParaGator-Zero-4B-Principia + aggregationBackbone=Qwen3-4B-Base, Aggregation=Enabled2026.03 | 27.4 | 29.24 | 45.17 | 33.24 | 33.77 | |
| Polaris-4BTraining Data=Polaris-Data.2026.03 | 26.17 | 51.02 | 64.36 | 45.82 | 46.84 | |
| Offline aggregation (AggLM)Backbone=Qwen3-4B-Instruct-2507, Aggregation=None2026.03 | 25.19 | 40.05 | 62.75 | 44.46 | 43.11 | |
| Qwen3-4B (thinking)Training Data=-2026.03 | 23.81 | 40.57 | 58.78 | 41.77 | 41.23 | |
| Online Multitask + aggregationBackbone=Qwen3-4B-Base, Aggregation=Enabled2026.03 | 23.07 | 25.31 | 43.4 | 29.65 | 30.36 | |
| Qwen3-4B-Instruct-2507Backbone=Qwen3-4B-Instruct-2507, Aggregation=None2026.03 | 22.94 | 33.75 | 63.3 | 41.87 | 40.47 | |
| Qwen3-14B (no-thinking)Training Data=-2026.03 | 21.34 | 28.64 | 50.27 | 36.5 | 34.19 | |
| Dr.GRPO + aggregationBackbone=Qwen3-4B-Base, Aggregation=Enabled2026.03 | 19.32 | 21.06 | 44.17 | 32.28 | 29.21 | |
| Principia-4B-ZeroTraining Data=Principia Collec.2026.03 | 19.28 | 21.81 | 43.62 | 33.92 | 29.66 | |
| Llama-3.3-70B-InstructTraining Data=-2026.03 | 18.41 | 21.36 | 37.5 | 25.81 | 25.77 | |
| Online MultitaskBackbone=Qwen3-4B-Base, Aggregation=None2026.03 | 18.05 | 16.33 | 33.26 | 24.18 | 22.96 | |
| Qwen3-4B (no-thinking)Training Data=-2026.03 | 17.86 | 22.39 | 39.89 | 28.93 | 27.27 | |
| ParaGator-Zero-4B-PrincipiaBackbone=Qwen3-4B-Base, Aggregation=None2026.03 | 17.71 | 21.09 | 38.62 | 28.75 | 26.54 | |
| Dr.GRPOBackbone=Qwen3-4B-Base, Aggregation=None2026.03 | 17.4 | 15.65 | 35.26 | 29.3 | 24.4 | |
| General-Reasoner-4BTraining Data=WebInstruct-Ver.2026.03 | 16.06 | 18.07 | 39.36 | 27.88 | 25.34 | |
| General-Reasoner-7BTraining Data=WebInstruct-Ver.2026.03 | 15.96 | 12.39 | 26.86 | 23.15 | 19.59 | |
| Principia-7B-ZeroTraining Data=Principia Collec.2026.03 | 15.59 | 15.11 | 32.45 | 28.34 | 22.87 | |
| OpenReasoner-ZeroTraining Data=ORZ-Math-Coll.2026.03 | 15.09 | 13.75 | 30.85 | 25.12 | 21.2 | |
| SimpleRL-7B-ZooTraining Data=SimpleZoo-Data.2026.03 | 14 | 10.68 | 26.86 | 21.17 | 18.18 | |
| Princpia-8B-ZeroTraining Data=Principia Collec.2026.03 | 13.57 | 14.2 | 28.46 | 19.96 | 19.05 | |
| Qwen2.5-7B-InstructTraining Data=-2026.03 | 12.95 | 10.45 | 19.15 | 20.05 | 15.65 | |
| Offline aggregation (AggLM) + aggregationBackbone=Qwen3-4B-Base, Aggregation=Enabled2026.03 | 12.1 | 19.64 | 31.27 | 21.62 | 21.15 | |
| Qwen2.5-7B-BaseTraining Data=-2026.03 | 11.19 | 9.32 | 16.76 | 13.75 | 12.75 | |
| DeepScaleRBase Model=OctoThinker-8B-Long-Base, Status=Reimplemented, Training Data=DeepScaleR2026.03 | 10.66 | 11.02 | 19.95 | 16.82 | 14.61 | |
| WebInstruct-Ver.Base Model=OctoThinker-8B-Long-Base, Status=Reimplemented, Training Data=WebInstruct-Ver.2026.03 | 10.66 | 11.02 | 19.95 | 20.56 | 17.67 | |
| Offline aggregation (AggLM)Backbone=Qwen3-4B-Base, Aggregation=None2026.03 | 10.15 | 12.37 | 25.4 | 15.26 | 15.8 | |
| Qwen3-4B-BaseTraining Data=-2026.03 | 9.43 | 5.8 | 17.81 | 12.18 | 11.31 | |
| DeepScaleRBase Model=Qwen3-4B-Base, Status=Reimplemented, Training Data=DeepScaleR2026.03 | 9.24 | 20.91 | 38.3 | 31.04 | 27.42 | |
| Qwen3-4B-BaseBackbone=Qwen3-4B-Base, Aggregation=None2026.03 | 9.11 | 5.69 | 17.44 | 11.01 | 10.81 | |
| Qwen3-4B-Base + aggregationBackbone=Qwen3-4B-Base, Aggregation=Enabled2026.03 | 8.35 | 4.79 | 18.24 | 10.68 | 10.52 | |
| Llama-3.1-8B-InstructTraining Data=-2026.03 | 6.01 | 5.8 | 9.31 | 7.4 | 7.13 | |
| Llama-3.2-3B-InstructTraining Data=-2026.03 | 3.7 | 3.3 | 1.33 | 4.25 | 3.14 | |
| OctoThinker-8B-Long-BaseTraining Data=-2026.03 | 3.16 | 2.73 | 5.32 | 3.79 | 3.75 |