Scientific Reasoning on GPQA Diamond (Accuracy, Avg.)
74.24AccuracyMILES
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| MILESBackbone=GPT-OSS-20B, Inference Strategy=MILES2026.07 | 74.24 | — | |
| MILESBackbone=Qwen3-30B-Instruct, Inference Strategy=MILES2026.07 | 73.74 | — | |
| MILESBackbone=GPT-4.1-mini, Inference Strategy=MILES2026.07 | 73.23 | — | |
| SCBackbone=Qwen3-30B-Instruct, Inference Strategy=SC2026.07 | 72.23 | — | |
| SCBackbone=GPT-4.1-mini, Inference Strategy=SC2026.07 | 71.72 | — | |
| SCBackbone=GPT-OSS-20B, Inference Strategy=SC2026.07 | 71.72 | — | |
| MILESBackbone=GPT-4.1, Inference Strategy=MILES2026.07 | 70.71 | — | |
| ZS-CoTBackbone=GPT-4.1-mini, Inference Strategy=ZS-CoT2026.07 | 70.2 | — | |
| ZS-CoTBackbone=Qwen3-30B-Instruct, Inference Strategy=ZS-CoT2026.07 | 70.2 | — | |
| ZS-CoTBackbone=GPT-4.1, Inference Strategy=ZS-CoT2026.07 | 68.94 | — | |
| SCBackbone=GPT-4.1, Inference Strategy=SC2026.07 | 68.69 | — | |
| ZS-CoTBackbone=GPT-OSS-20B, Inference Strategy=ZS-CoT2026.07 | 68.18 | — | |
| BoTBackbone=Qwen3-30B-Instruct, Inference Strategy=BoT2026.07 | 67.68 | — | |
| BoTBackbone=GPT-4.1-mini, Inference Strategy=BoT2026.07 | 66.67 | — | |
| DCBackbone=GPT-4.1, Inference Strategy=DC2026.07 | 64.65 | — | |
| DCBackbone=GPT-4.1-mini, Inference Strategy=DC2026.07 | 63.64 | — | |
| DCBackbone=GPT-OSS-20B, Inference Strategy=DC2026.07 | 61.62 | — | |
| BoTBackbone=GPT-4.1, Inference Strategy=BoT2026.07 | 61.11 | — | |
| TeacherModel=Qwen3-4B2026.07 | 59 | — | |
| KVPOP_mlpModel=Qwen3-4B, Compression Ratio=75%2026.07 | 59 | — | |
| TOVAModel=Qwen3-4B, Compression Ratio=75%2026.07 | 58 | — | |
| KVPOPModel=Qwen3-4B, Compression Ratio=75%2026.07 | 57 | — | |
| KVPOP_mlpModel=Qwen3-4B, Compression Ratio=88%2026.07 | 57 | — | |
| KVPOPModel=Qwen3-4B, Compression Ratio=88%2026.07 | 56 | — | |
| StreamLLM+Model=Qwen3-4B, Compression Ratio=75%2026.07 | 55 | — | |
| DMSModel=Qwen3-4B, Compression Ratio=75%2026.07 | 55 | — | |
| StreamLLMModel=Qwen3-4B, Compression Ratio=75%2026.07 | 54 | — | |
| TOVAModel=Qwen3-4B, Compression Ratio=88%2026.07 | 54 | — | |
| DMSModel=Qwen3-4B, Compression Ratio=88%2026.07 | 54 | — | |
| StreamLLM+Model=Qwen3-4B, Compression Ratio=88%2026.07 | 52 | — | |
| Random BaselineBase Model=OpenR1-Distill-7B2026.06 | 51.45 | — | |
| DRIFTBase Model=OpenR1-Distill-7B2026.06 | 50.54 | — | |
| RDSBase Model=OpenR1-Distill-7B2026.06 | 50.19 | — | |
| OpenR1-Distill-7BBase Model=OpenR1-Distill-7B2026.06 | 49.72 | — | |
| BM25Base Model=OpenR1-Distill-7B2026.06 | 49.4 | — | |
| IF (Impl. w. GraSS)Base Model=OpenR1-Distill-7B2026.06 | 49.15 | — | |
| DSIRBase Model=OpenR1-Distill-7B2026.06 | 49.02 | — | |
| LESSBase Model=OpenR1-Distill-7B2026.06 | 49.02 | — | |
| StreamLLMModel=Qwen3-4B, Compression Ratio=88%2026.07 | 49 | — | |
| QuratingBase Model=OpenR1-Distill-7B2026.06 | 47.57 | — | |
| Self-DistillationBase Model=OpenR1-Distill-7B2026.06 | 44.48 | — | |
| DCBackbone=Qwen3-30B-Instruct, Inference Strategy=DC2026.07 | 42.42 | — | |
| PolicyAlignBackbone=Qwen2.5-14B-Instruct, Alignment Strategy=Policy-Based Safety Alignment2026.06 | 38.38 | — | |
| BaseBackbone=Qwen2.5-14B-Instruct, Alignment Strategy=Base2026.06 | 37.88 | — | |
| GRPO+PolicyBackbone=Qwen2.5-14B-Instruct, Alignment Strategy=Group Relative Policy Optimization with policy-conditioned reward2026.06 | 37.37 | — | |
| BoTBackbone=GPT-OSS-20B, Inference Strategy=BoT2026.07 | 37.37 | — | |
| NSPOBackbone=Qwen2.5-14B-Instruct, Alignment Strategy=Null-space constrained policy optimization2026.06 | 36.87 | — | |
| CoTBackbone=Qwen2.5-7B-Instruct, Method Category=Non-RL Baselines2026.05 | 36.4 | 67.83 | |
| SFTBackbone=Qwen2.5-14B-Instruct, Alignment Strategy=Supervised Fine-Tuning2026.06 | 36.36 | — | |
| BM25Base Model=Olmo3-7B-Instruct-SFT2026.06 | 35.92 | — | |
| AlphaAlignBackbone=Qwen2.5-14B-Instruct, Alignment Strategy=RL-based alignment2026.06 | 35.86 | — | |
| TRACERBackbone=Qwen2.5-7B-Instruct, Method Category=Non-RL/RL Baseline Representative2026.05 | 35.35 | 61.85 | |
| QuratingBase Model=Olmo3-7B-Instruct-SFT2026.06 | 35.23 | — | |
| DRIFTBase Model=Olmo3-7B-Instruct-SFT2026.06 | 35.23 | — | |
| Olmo3-7B-Instruct-SFTBase Model=Olmo3-7B-Instruct-SFT2026.06 | 35.01 | — | |
| RDSBase Model=Olmo3-7B-Instruct-SFT2026.06 | 34.5 | — | |
| IF (Impl. w. GraSS)Base Model=Olmo3-7B-Instruct-SFT2026.06 | 34.44 | — | |
| ICLBackbone=Qwen2.5-14B-Instruct, Alignment Strategy=In-Context Learning (System Prompt)2026.06 | 33.84 | — | |
| DSIRBase Model=Olmo3-7B-Instruct-SFT2026.06 | 33.62 | — | |
| Random BaselineBase Model=Olmo3-7B-Instruct-SFT2026.06 | 33.14 | — | |
| CoTBackbone=Phi-3 Mini 4K Instruct, Method Category=Non-RL Baselines2026.05 | 32.8 | 52.2 | |
| LESSBase Model=Olmo3-7B-Instruct-SFT2026.06 | 32.8 | — | |
| MADBackbone=Qwen2.5-7B-Instruct, Method Category=Non-RL Baselines2026.05 | 32.1 | 56.11 | |
| MADBackbone=Phi-3 Mini 4K Instruct, Method Category=Non-RL Baselines2026.05 | 31.82 | 51.1 | |
| Sparse MADBackbone=Phi-3 Mini 4K Instruct, Method Category=Non-RL Baselines2026.05 | 31.31 | 51.13 | |
| Self-DistillationBase Model=Olmo3-7B-Instruct-SFT2026.06 | 30.93 | — | |
| Sparse MADBackbone=Qwen2.5-7B-Instruct, Method Category=Non-RL Baselines2026.05 | 30.81 | 55.8 | |
| Self-ConsistencyBackbone=Phi-3 Mini 4K Instruct, Method Category=Non-RL Baselines2026.05 | 30.8 | 52.32 | |
| ICLBackbone=Qwen2.5-7B-Instruct, Alignment Strategy=In-Context Learning (System Prompt)2026.06 | 28.78 | — | |
| MoABackbone=Qwen2.5-7B-Instruct, Method Category=Non-RL Baselines2026.05 | 28.28 | 53.39 | |
| PolicyAlignBackbone=Qwen2.5-7B-Instruct, Alignment Strategy=Policy-Based Safety Alignment2026.06 | 28.28 | — | |
| SFTBackbone=Qwen2.5-7B-Instruct, Alignment Strategy=Supervised Fine-Tuning2026.06 | 27.78 | — | |
| GRPO+PolicyBackbone=Qwen2.5-7B-Instruct, Alignment Strategy=Group Relative Policy Optimization with policy-conditioned reward2026.06 | 27.78 | — | |
| Self-ConsistencyBackbone=Qwen2.5-7B-Instruct, Method Category=Non-RL Baselines2026.05 | 27.27 | 62.42 | |
| TRACERBackbone=Phi-3 Mini 4K Instruct, Method Category=Non-RL/RL Baseline Representative2026.05 | 27.27 | 50.27 | |
| AlphaAlignBackbone=Qwen2.5-7B-Instruct, Alignment Strategy=RL-based alignment2026.06 | 27.27 | — | |
| NSPOBackbone=Qwen2.5-7B-Instruct, Alignment Strategy=Null-space constrained policy optimization2026.06 | 27.27 | — | |
| MAPoRLBackbone=Phi-3 Mini 4K Instruct, Method Category=RL Baselines2026.05 | 26.64 | 48.93 | |
| BaseBackbone=Qwen2.5-7B-Instruct, Alignment Strategy=Base2026.06 | 26.26 | — | |
| MoABackbone=Phi-3 Mini 4K Instruct, Method Category=Non-RL Baselines2026.05 | 24.75 | 47.26 | |
| PolicyAlignBackbone=LLaMA-3.2-3B-Instruct, Alignment Strategy=Policy-Based Safety Alignment2026.06 | 24.24 | — | |
| BaseBackbone=LLaMA-3.2-3B-Instruct, Alignment Strategy=Base2026.06 | 23.74 | — | |
| NSPOBackbone=LLaMA-3.2-3B-Instruct, Alignment Strategy=Null-space constrained policy optimization2026.06 | 23.23 | — | |
| GRPO+PolicyBackbone=LLaMA-3.2-3B-Instruct, Alignment Strategy=Group Relative Policy Optimization with policy-conditioned reward2026.06 | 23.23 | — | |
| ICLBackbone=LLaMA-3.2-3B-Instruct, Alignment Strategy=In-Context Learning (System Prompt)2026.06 | 22.22 | — | |
| AlphaAlignBackbone=LLaMA-3.2-3B-Instruct, Alignment Strategy=RL-based alignment2026.06 | 22.22 | — | |
| SFTBackbone=LLaMA-3.2-3B-Instruct, Alignment Strategy=Supervised Fine-Tuning2026.06 | 21.72 | — | |
| MAPoRLBackbone=Qwen2.5-7B-Instruct, Method Category=RL Baselines2026.05 | 20.71 | 38.19 | |
| Single-Agent GRPOBackbone=Phi-3 Mini 4K Instruct, Method Category=RL Baselines2026.05 | 9.47 | 46.63 | |
| MAGRPOBackbone=Phi-3 Mini 4K Instruct, Method Category=RL Baselines2026.05 | 8.84 | 38.52 | |
| Single-Agent GSPOBackbone=Phi-3 Mini 4K Instruct, Method Category=RL Baselines2026.05 | 8.21 | 45.72 | |
| Single-Agent GSPOBackbone=Qwen2.5-7B-Instruct, Method Category=RL Baselines2026.05 | 3.66 | 45.63 | |
| Single-Agent GRPOBackbone=Qwen2.5-7B-Instruct, Method Category=RL Baselines2026.05 | 3.41 | 46.49 | |
| MAGRPOBackbone=Qwen2.5-7B-Instruct, Method Category=RL Baselines2026.05 | 1.89 | 34.73 | |
| Base (SFT)inference_mode=thinking mode2026.06 | — | 64.65 | |
| Claude Opus 4.82026.06 | — | 92 | |
| Fugu2026.06 | — | 95.5 | |
| Fugu-Ultra2026.06 | — | 95.5 | |
| Gemini 3.12026.06 | — | 94.3 | |
| GPT-5.52026.06 | — | 93.6 |