Scientific Reasoning on GPQA Diamond (Pass@1, Token count)
78.66Pass@1 AccuracyGPT-5.2-chat (teacher)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-5.2-chat (teacher)2026.05 | 78.66 | — | |
| AlphaOneModel=Qwen3-32B-thinking2026.04 | 66.8 | 5,591 | |
| REDModel=Qwen3-32B-thinking2026.04 | 66.8 | 2,477 | |
| S-GRPOModel=Qwen3-32B-thinking2026.04 | 66.5 | 4,151 | |
| GRPOModel=Qwen3-32B-thinking2026.04 | 66.2 | 6,151 | |
| REDModel=DeepSeek-R1-Distill-Llama-70B2026.04 | 66.2 | 3,563 | |
| RL + LPModel=Qwen3-32B-thinking2026.04 | 66 | 4,362 | |
| DASTModel=Qwen3-32B-thinking2026.04 | 65.7 | 2,918 | |
| VanillaModel=Qwen3-32B-thinking2026.04 | 65.3 | 5,475 | |
| AlphaOneModel=DeepSeek-R1-Distill-Llama-70B2026.04 | 65.3 | 4,322 | |
| S-GRPOModel=DeepSeek-R1-Distill-Llama-70B2026.04 | 65.2 | 4,001 | |
| RL + LPModel=DeepSeek-R1-Distill-Llama-70B2026.04 | 65 | 4,101 | |
| DEERModel=Qwen3-32B-thinking2026.04 | 64.8 | 2,247 | |
| GRPOModel=DeepSeek-R1-Distill-Llama-70B2026.04 | 64.8 | 6,274 | |
| Think or NotModel=DeepSeek-R1-Distill-Llama-70B2026.04 | 64.6 | 3,544 | |
| Think or NotModel=Qwen3-32B-thinking2026.04 | 64.5 | 1,713 | |
| VanillaModel=DeepSeek-R1-Distill-Llama-70B2026.04 | 64.5 | 5,881 | |
| DASTModel=DeepSeek-R1-Distill-Llama-70B2026.04 | 63.5 | 4,026 | |
| DEERModel=DeepSeek-R1-Distill-Llama-70B2026.04 | 63.1 | 4,502 | |
| REDModel=DpSk-R1-Distill-Qwen-32B2026.04 | 61.8 | 3,497 | |
| AlphaOneModel=DpSk-R1-Distill-Qwen-32B2026.04 | 61.3 | 6,771 | |
| S-GRPOModel=DpSk-R1-Distill-Qwen-32B2026.04 | 61.3 | 3,119 | |
| VanillaModel=DpSk-R1-Distill-Qwen-32B2026.04 | 60.8 | 6,027 | |
| Think or NotModel=DpSk-R1-Distill-Qwen-32B2026.04 | 60.4 | 3,706 | |
| RL + LPModel=DpSk-R1-Distill-Qwen-32B2026.04 | 60.4 | 3,596 | |
| GRPOModel=DpSk-R1-Distill-Qwen-32B2026.04 | 60.4 | 7,123 | |
| DASTModel=DpSk-R1-Distill-Qwen-32B2026.04 | 60.3 | 4,048 | |
| REDModel=Qwen3-8B-thinking2026.04 | 60.1 | 3,309 | |
| AlphaOneModel=Qwen3-8B-thinking2026.04 | 59.8 | 6,002 | |
| S-GRPOModel=Qwen3-8B-thinking2026.04 | 59.6 | 3,046 | |
| RL + LPModel=Qwen3-8B-thinking2026.04 | 59.4 | 3,139 | |
| Think or NotModel=Qwen3-8B-thinking2026.04 | 59.1 | 3,678 | |
| GRPOModel=Qwen3-8B-thinking2026.04 | 59.1 | 7,472 | |
| DEERModel=DpSk-R1-Distill-Qwen-32B2026.04 | 58.9 | 4,611 | |
| VanillaModel=Qwen3-8B-thinking2026.04 | 58.8 | 6,638 | |
| DASTModel=Qwen3-8B-thinking2026.04 | 58.4 | 3,007 | |
| DEERModel=Qwen3-8B-thinking2026.04 | 56.9 | 2,926 | |
| ROPDThinking Mode=Thinking2026.05 | 55.05 | — | |
| OVDThinking Mode=Thinking2026.05 | 54.17 | — | |
| T-JudgeThinking Mode=Thinking2026.05 | 53.85 | — | |
| GADThinking Mode=Thinking2026.05 | 53.85 | — | |
| MAPRBase Model=Qwen3-14B-Base2025.09 | 53.72 | — | |
| Qwen3-4B (student)Thinking Mode=Thinking2026.05 | 53.59 | — | |
| GRPOBase Model=Qwen3-14B-Base2025.09 | 51.72 | — | |
| REDModel=DpSk-R1-Distill-Qwen-7B2026.04 | 51.2 | 4,109 | |
| AlphaOneModel=DpSk-R1-Distill-Qwen-7B2026.04 | 50.5 | 8,591 | |
| RL + LPModel=DpSk-R1-Distill-Qwen-7B2026.04 | 50.3 | 3,209 | |
| GRPOModel=DpSk-R1-Distill-Qwen-7B2026.04 | 49.8 | 8,890 | |
| S-GRPOModel=DpSk-R1-Distill-Qwen-7B2026.04 | 49.5 | 3,107 | |
| VanillaModel=DpSk-R1-Distill-Qwen-7B2026.04 | 49.2 | 8,016 | |
| DASTModel=DpSk-R1-Distill-Qwen-7B2026.04 | 48.8 | 3,635 | |
| REDModel=DeepSeek-R1-Distill-Llama-8B2026.04 | 47.8 | 3,593 | |
| AlphaOneModel=DeepSeek-R1-Distill-Llama-8B2026.04 | 47.6 | 8,569 | |
| DEERModel=DpSk-R1-Distill-Qwen-7B2026.04 | 47.3 | 4,423 | |
| Think or NotModel=DpSk-R1-Distill-Qwen-7B2026.04 | 47 | 3,390 | |
| S-GRPOModel=DeepSeek-R1-Distill-Llama-8B2026.04 | 47 | 3,624 | |
| Think or NotModel=DeepSeek-R1-Distill-Llama-8B2026.04 | 46.8 | 3,729 | |
| GRPOModel=DeepSeek-R1-Distill-Llama-8B2026.04 | 46.6 | 8,783 | |
| VanillaModel=DeepSeek-R1-Distill-Llama-8B2026.04 | 46.3 | 8,341 | |
| DASTModel=DeepSeek-R1-Distill-Llama-8B2026.04 | 46.1 | 4,410 | |
| DEERModel=DeepSeek-R1-Distill-Llama-8B2026.04 | 45.5 | 4,152 | |
| RL + LPModel=DeepSeek-R1-Distill-Llama-8B2026.04 | 45.3 | 3,299 | |
| ROPDThinking Mode=Non-Thinking2026.05 | 36.5 | — | |
| T-JudgeThinking Mode=Non-Thinking2026.05 | 36.29 | — | |
| GADThinking Mode=Non-Thinking2026.05 | 36.02 | — | |
| OVDThinking Mode=Non-Thinking2026.05 | 35.74 | — | |
| Qwen3-4B (student)Thinking Mode=Non-Thinking2026.05 | 35.66 | — |