Scientific Reasoning on GPQA Diamond (Accuracy)
80.13AccuracyUpper Bound
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Upper BoundModel=Qwen3-8B, k=62025.08 | 80.13 | — | |
| Upper BoundModel=DS-Distill-Qwen-2.5-7B, k=62025.08 | 79.12 | — | |
| Upper BoundModel=DS-Distill-llama-3-8B, k=62025.08 | 77.27 | — | |
| QwenLong-L1.5-30B-A3B2025.12 | 76.78 | — | |
| Qwen3-30B-A3B-Thinking-25072025.12 | 75.88 | — | |
| LamPOBackbone=Qwen3-4B2026.05 | 69.42 | — | |
| GRPOBackbone=Qwen3-4B2026.05 | 67.74 | — | |
| CoTBackbone=Qwen3-4B2026.05 | 67.17 | — | |
| Teacher-30BInference Mode=thinking2026.06 | 66.67 | — | |
| Critique-GRPOCritique Model=GPT-4o, Decoding budget=16,384 tokens2025.06 | 60.61 | — | |
| Teacher-8BInference Mode=thinking2026.06 | 59.6 | — | |
| PiCSARModel=Qwen3-8B, k=62025.08 | 59.43 | — | |
| Base (SFT ckpt)Inference Mode=thinking2026.06 | 57.58 | — | |
| RLInference Mode=thinking2026.06 | 57.58 | — | |
| R1-GRPOCritique Model=None, Decoding budget=16,384 tokens2025.06 | 56.57 | — | |
| MiMo-V2-Flash Base# Shots=5-shot, # Activated Params=15B, # Total Params=309B2026.01 | 55.1 | — | |
| AverageModel=Qwen3-8B, k=62025.08 | 54.21 | — | |
| Self-ConsistencyModel=Qwen3-8B, k=62025.08 | 54.21 | — | |
| OPD-T8BInference Mode=thinking2026.06 | 54.05 | — | |
| OPD-T30BInference Mode=thinking2026.06 | 53.04 | — | |
| PiCSARModel=DS-Distill-Qwen-2.5-7B, k=62025.08 | 52.36 | — | |
| DeepSeek-V3.2 Exp Base# Shots=5-shot, # Activated Params=37B, # Total Params=671B2026.01 | 52 | — | |
| DeepSeek-V3.1 Base# Shots=5-shot, # Activated Params=37B, # Total Params=671B2026.01 | 51 | — | |
| LamPOBackbone=Phi-4-mini2026.05 | 48.44 | — | |
| Kimi-K2 Base# Shots=5-shot, # Activated Params=32B, # Total Params=1043B2026.01 | 48.1 | — | |
| PiCSARModel=DS-Distill-llama-3-8B, k=62025.08 | 47.31 | — | |
| Skywork-OR1-Math-7BDistillation Domain=Single-Domain, Teacher Model=Skywork-OR1-Math-7B2026.05 | 47.22 | — | |
| Skywork-OR1-7BDistillation Domain=Multi-Domain, Teacher Model=Skywork-OR1-7B2026.05 | 47.22 | — | |
| PiCSARModel=Qwen3-8B (Non-thinking), k=62025.08 | 46.98 | — | |
| PiCSARModel=Llama-3.1-70B-Instruct, k=62025.08 | 46.91 | — | |
| PiCSARModel=Qwen3-32B (Non-thinking), k=62025.08 | 46.91 | — | |
| GRPOBackbone=Phi-4-mini2026.05 | 46.59 | — | |
| AverageModel=DS-Distill-Qwen-2.5-7B, k=62025.08 | 46.44 | — | |
| Qwen3-32BCritique Model=None, Decoding budget=16,384 tokens2025.06 | 45.96 | — | |
| CoTBackbone=Phi-4-mini2026.05 | 45.33 | — | |
| Self-ConsistencyModel=DS-Distill-Qwen-2.5-7B, k=62025.08 | 44.78 | — | |
| DYPOBase Model=Qwen3-4B-Base2026.04 | 44.4 | — | |
| GSPOBackbone=Qwen3-1.7B2026.05 | 42.91 | — | |
| AverageModel=DS-Distill-llama-3-8B, k=62025.08 | 42.87 | — | |
| LamPOBackbone=Qwen3-1.7B2026.05 | 42.86 | — | |
| RLPRBackbone=Qwen3-1.7B2026.05 | 42.42 | — | |
| Self-ConsistencyModel=DS-Distill-llama-3-8B, k=62025.08 | 42.1 | — | |
| PRIMEBackbone=Qwen3-1.7B2026.05 | 41.89 | — | |
| GRPOBackbone=Qwen3-1.7B2026.05 | 41.78 | — | |
| SimpleRL-ZooBackbone=Qwen3-1.7B2026.05 | 41.41 | — | |
| DAPOBackbone=Qwen3-1.7B2026.05 | 41.4 | — | |
| CoTBackbone=Qwen3-1.7B2026.05 | 38.78 | — | |
| TrOPDDistillation Domain=Single-Domain, Teacher Model=Skywork-OR1-Math-7B, Student Model=DeepSeek-R1-Distill-Qwen-1.5B, Training Steps=200, Learning Rate=5 x 10^-6, Sample Rollouts=4, Max Generation Length=80962026.05 | 36.24 | — | |
| REOPOLDDistillation Domain=Multi-Domain, Teacher Model=Skywork-OR1-7B, Student Model=DeepSeek-R1-Distill-Qwen-1.5B, Training Steps=200, Learning Rate=5 x 10^-6, Sample Rollouts=4, Max Generation Length=80962026.05 | 34.47 | — | |
| DeepSeek-Qwen2.5-1.5BStudent Model=DeepSeek-Qwen2.5-1.5B2026.05 | 34.22 | — | |
| DAPOBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.04 | 33.33 | — | |
| EOPDDistillation Domain=Single-Domain, Teacher Model=Skywork-OR1-Math-7B, Student Model=DeepSeek-R1-Distill-Qwen-1.5B, Training Steps=200, Learning Rate=5 x 10^-6, Sample Rollouts=4, Max Generation Length=80962026.05 | 32.58 | — | |
| PiCSARModel=Gemma-2-9B-Instruct, k=62025.08 | 32.32 | — | |
| REOPOLDDistillation Domain=Single-Domain, Teacher Model=Skywork-OR1-Math-7B, Student Model=DeepSeek-R1-Distill-Qwen-1.5B, Training Steps=200, Learning Rate=5 x 10^-6, Sample Rollouts=4, Max Generation Length=80962026.05 | 32.07 | — | |
| Entropy OPD 20%Distillation Domain=Single-Domain, Teacher Model=Skywork-OR1-Math-7B, Student Model=DeepSeek-R1-Distill-Qwen-1.5B, Training Steps=200, Learning Rate=5 x 10^-6, Sample Rollouts=4, Max Generation Length=80962026.05 | 31.82 | — | |
| TrOPDDistillation Domain=Multi-Domain, Teacher Model=Skywork-OR1-7B, Student Model=DeepSeek-R1-Distill-Qwen-1.5B, Training Steps=200, Learning Rate=5 x 10^-6, Sample Rollouts=4, Max Generation Length=80962026.05 | 31.19 | — | |
| OPDDistillation Domain=Multi-Domain, Teacher Model=Skywork-OR1-7B, Student Model=DeepSeek-R1-Distill-Qwen-1.5B, Training Steps=200, Learning Rate=5 x 10^-6, Sample Rollouts=4, Max Generation Length=80962026.05 | 31.06 | — | |
| REOPOLD 2StageDistillation Domain=Single-Domain, Teacher Model=Skywork-OR1-Math-7B, Student Model=DeepSeek-R1-Distill-Qwen-1.5B, Training Steps=200, Learning Rate=5 x 10^-6, Sample Rollouts=4, Max Generation Length=80962026.05 | 30.18 | — | |
| SPSBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.04 | 29.8 | — | |
| PiCSARModel=Llama-3.1-8B-Instruct, k=62025.08 | 29.8 | — | |
| SDAR-1.7BScale=1.7B, Tokens=55B, FLOPs (1e18)=561.02026.06 | 29.8 | — | |
| RLBase Model=Qwen3-4B-Base2026.04 | 29.3 | — | |
| SFTBase Model=Qwen3-4B-Base2026.04 | 28.8 | — | |
| OPDDistillation Domain=Single-Domain, Teacher Model=Skywork-OR1-Math-7B, Student Model=DeepSeek-R1-Distill-Qwen-1.5B, Training Steps=200, Learning Rate=5 x 10^-6, Sample Rollouts=4, Max Generation Length=80962026.05 | 28.03 | — | |
| SFT -> RLBase Model=Qwen3-4B-Base2026.04 | 27.8 | — | |
| OPDLM-1.7BScale=1.7B, Tokens=0.072B, FLOPs (1e18)=0.982026.06 | 26.8 | — | |
| Fast-dLLM-v2-1.5BScale=1.5B, Tokens=1B, FLOPs (1e18)=9.02026.06 | 24.8 | — | |
| OPDLM-0.6BScale=0.6B, Tokens=0.048B, FLOPs (1e18)=0.232026.06 | 24 | — | |
| Simple-dLLMScale=0.6B, Tokens=46B, FLOPs (1e18)=165.62026.06 | 22.2 | — | |
| GSPOBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.04 | 21.21 | — | |
| GRPOBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.04 | 19.7 | — | |
| DeepSeek-R1-Distill-Qwen-1.5BType=Base Model2026.04 | 14.14 | — | |
| Qwen3-4B-BaseBase Model=Qwen3-4B-Base2026.04 | 14.1 | — | |
| BF16Model=DeepSeek-R1-Distill-Llama-8B, Bit-width=BF162025.12 | — | 48.99 | |
| BF16Model=DeepSeek-R1-Distill-Qwen-14B, Bit-width=BF162025.12 | — | 58.59 | |
| BF16Model=DeepSeek-R1-Distill-Qwen-32B, Bit-width=BF162025.12 | — | 62.63 | |
| KIVIModel=DeepSeek-R1-Distill-Llama-8B, Bit-width=KV42025.12 | — | 46.46 | |
| KIVIModel=DeepSeek-R1-Distill-Llama-8B, Bit-width=KV22025.12 | — | 37.37 | |
| KIVIModel=DeepSeek-R1-Distill-Qwen-14B, Bit-width=KV42025.12 | — | 56.06 | |
| KIVIModel=DeepSeek-R1-Distill-Qwen-14B, Bit-width=KV22025.12 | — | 52.53 | |
| KIVIModel=DeepSeek-R1-Distill-Qwen-32B, Bit-width=KV42025.12 | — | 59.09 | |
| KIVIModel=DeepSeek-R1-Distill-Qwen-32B, Bit-width=KV22025.12 | — | 54.54 | |
| KVQuantModel=DeepSeek-R1-Distill-Llama-8B, Bit-width=KV42025.12 | — | 44.95 | |
| KVQuantModel=DeepSeek-R1-Distill-Llama-8B, Bit-width=KV22025.12 | — | 21.72 | |
| KVQuantModel=DeepSeek-R1-Distill-Qwen-14B, Bit-width=KV42025.12 | — | 54.58 | |
| KVQuantModel=DeepSeek-R1-Distill-Qwen-14B, Bit-width=KV22025.12 | — | 22.22 | |
| KVQuantModel=DeepSeek-R1-Distill-Qwen-32B, Bit-width=KV42025.12 | — | 58.08 | |
| KVQuantModel=DeepSeek-R1-Distill-Qwen-32B, Bit-width=KV22025.12 | — | 30.81 | |
| KVTunerModel=DeepSeek-R1-Distill-Llama-8B, Bit-width=C3.252025.12 | — | 46.96 | |
| KVTunerModel=DeepSeek-R1-Distill-Qwen-14B, Bit-width=C2.902025.12 | — | 50 | |
| KVTunerModel=DeepSeek-R1-Distill-Qwen-32B, Bit-width=C2.912025.12 | — | 59.6 | |
| Mathstral-7B-v0.1Base Model=Mathstral-7B-v0.12025.09 | — | 29.3 | |
| Mathstral-7B-v0.1 + GRPOBase Model=Mathstral-7B-v0.1, Training=GRPO2025.09 | — | 28.8 | |
| Mathstral-7B-v0.1 + GRPO w/ VERL.Base Model=Mathstral-7B-v0.1, Training=GRPO w/ VERL.2025.09 | — | 31.8 | |
| Mathstral-7B-v0.1 + PPOBase Model=Mathstral-7B-v0.1, Training=PPO2025.09 | — | 28.8 | |
| Mathstral-7B-v0.1 + PPO w/ VERL.Base Model=Mathstral-7B-v0.1, Training=PPO w/ VERL.2025.09 | — | 31.3 | |
| MixKVQModel=DeepSeek-R1-Distill-Llama-8B, Bit-width=C2.72025.12 | — | 47.47 | |
| MixKVQModel=DeepSeek-R1-Distill-Qwen-14B, Bit-width=C2.32025.12 | — | 56.57 | |
| MixKVQModel=DeepSeek-R1-Distill-Qwen-32B, Bit-width=C2.32025.12 | — | 60.1 | |
| Qwen2.5-7BBase Model=Qwen2.5-7B2025.09 | — | 27.3 |