Scientific Reasoning on GPQA Diamond (Accuracy, Pass@1)
87.5AccuracyConductor
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Conductorparameters=7B2025.12 | 87.5 | — | |
| Gemini 2.5 Pro2025.12 | 84.8 | — | |
| GPT 52025.12 | 82.3 | — | |
| gpt-oss-120BReasoning effort=High, KV cache precision=BF162026.02 | 77.78 | — | |
| Claude Sonnet 42025.12 | 77.7 | — | |
| gpt-oss-120BReasoning effort=High, KV cache precision=FP82026.02 | 77.34 | — | |
| gpt-oss-puzzle-88BReasoning effort=High, KV cache precision=FP82026.02 | 75.25 | — | |
| gpt-oss-puzzle-88BReasoning effort=High, KV cache precision=BF162026.02 | 75.13 | — | |
| gpt-oss-120BReasoning effort=Medium, KV cache precision=BF162026.02 | 71.28 | — | |
| gpt-oss-puzzle-88BReasoning effort=Medium, KV cache precision=FP82026.02 | 71.15 | — | |
| gpt-oss-puzzle-88BReasoning effort=Medium, KV cache precision=BF162026.02 | 69.7 | — | |
| gpt-oss-120BReasoning effort=Medium, KV cache precision=FP82026.02 | 68.88 | — | |
| Qwen3-32B (thinking)mode=thinking2025.12 | 66.8 | — | |
| gpt-oss-puzzle-88BReasoning effort=Low, KV cache precision=BF162026.02 | 64.77 | — | |
| Qwen3-32B2025.12 | 64.1 | — | |
| gpt-oss-puzzle-88BReasoning effort=Low, KV cache precision=FP82026.02 | 63.51 | — | |
| gpt-oss-120BReasoning effort=Low, KV cache precision=BF162026.02 | 62.75 | — | |
| gpt-oss-120BReasoning effort=Low, KV cache precision=FP82026.02 | 60.54 | — | |
| R1-Distill-Qwen-32B2025.12 | 58.1 | — | |
| BF16Model=Qwen3-4B, W-Bits=BF162026.01 | 56.06 | — | |
| GPT-4on=198, Iterative Repair Setting=no iterative repair2026.04 | 49.4 | — | |
| Reasoning-QATModel=Qwen3-4B, W-Bits=W3G1282026.01 | 45.79 | — | |
| AeroTherm-GPTn=198, Iterative Repair Setting=no iterative repair2026.04 | 38.9 | — | |
| gemma-3-27b-itinstruct=true2025.12 | 38.4 | — | |
| GPTQModel=Qwen3-4B, W-Bits=W3G1282026.01 | 38.05 | — | |
| AWQModel=R1-Qwen-1.5B, W-Bits=W3G1282026.01 | 37.88 | — | |
| AWQModel=Qwen3-4B, W-Bits=W3G1282026.01 | 37.88 | — | |
| BF16Model=R1-Qwen-1.5B, W-Bits=BF162026.01 | 36.87 | — | |
| Llama-3-70Bn=198, Iterative Repair Setting=no iterative repair2026.04 | 33.1 | — | |
| Reasoning-QATModel=R1-Qwen-1.5B, W-Bits=W3G1282026.01 | 30.3 | — | |
| Segment Selective SFTModel Backbone=R1-Distill-Qwen-1.5B, Decoding Strategy=Greedy Decoding2026.01 | 28.8 | — | |
| BF16Model=Qwen3-0.6B, W-Bits=BF162026.01 | 28.45 | — | |
| Reasoning-QATModel=Qwen3-0.6B, W-Bits=W3G1282026.01 | 27.78 | — | |
| AWQModel=Qwen3-0.6B, W-Bits=W3G1282026.01 | 26.77 | — | |
| GPTQModel=Qwen3-0.6B, W-Bits=W3G1282026.01 | 26.43 | — | |
| Reasoning-QATModel=R1-Qwen-1.5B, W-Bits=W2G1282026.01 | 25.75 | — | |
| AWQModel=Qwen3-4B, W-Bits=W2G1282026.01 | 25.59 | — | |
| Reasoning-QATModel=Qwen3-4B, W-Bits=W2G1282026.01 | 25.42 | — | |
| Reasoning-QATModel=Qwen3-0.6B, W-Bits=W2G1282026.01 | 25.25 | — | |
| AWQModel=R1-Qwen-1.5B, W-Bits=W2G1282026.01 | 25.08 | — | |
| RTNModel=Qwen3-0.6B, W-Bits=W3G1282026.01 | 24.24 | — | |
| AWQModel=Qwen3-0.6B, W-Bits=W2G1282026.01 | 23.91 | — | |
| GPTQModel=R1-Qwen-1.5B, W-Bits=W3G1282026.01 | 23.74 | — | |
| GPTQModel=R1-Qwen-1.5B, W-Bits=W2G1282026.01 | 21.89 | — | |
| GPTQModel=Qwen3-4B, W-Bits=W2G1282026.01 | 20.7 | — | |
| RTNModel=R1-Qwen-1.5B, W-Bits=W3G1282026.01 | 19.19 | — | |
| RTNModel=Qwen3-4B, W-Bits=W3G1282026.01 | 10.6 | — | |
| GPTQModel=Qwen3-0.6B, W-Bits=W2G1282026.01 | 0.84 | — | |
| BaseBackbone=DeepSeek-R1-Distill-Llama-8B2025.09 | — | 44.9 | |
| BaseBackbone=DeepSeek-R1-Distill-Qwen-7B2025.09 | — | 47 | |
| BaseBackbone=Qwen3-8B2025.09 | — | 53.5 | |
| Critique-GRPO (CoT Critique)Backbone=Qwen2.5-7B-Base2025.06 | — | 37.88 | |
| Critique-GRPO (CoT Critique)Backbone=Qwen3-8B2025.06 | — | 47.98 | |
| GRPOBackbone=DeepSeek-R1-Distill-Llama-8B2025.09 | — | 50.5 | |
| GRPOBackbone=DeepSeek-R1-Distill-Qwen-7B2025.09 | — | 49 | |
| GRPOBackbone=Qwen3-8B2025.09 | — | 59.1 | |
| IPOBackbone=DeepSeek-R1-Distill-Llama-8B2025.09 | — | 49 | |
| IPOBackbone=DeepSeek-R1-Distill-Qwen-7B2025.09 | — | 51.5 | |
| IPOBackbone=Qwen3-8B2025.09 | — | 59.1 | |
| Qwen2.5-7B-BaseBackbone=Qwen2.5-7B-Base2025.06 | — | 28.79 | |
| Qwen3-8BBackbone=Qwen3-8B2025.06 | — | 35.86 | |
| RealSafeBackbone=DeepSeek-R1-Distill-Llama-8B2025.09 | — | 47.5 | |
| RealSafeBackbone=DeepSeek-R1-Distill-Qwen-7B2025.09 | — | 51 | |
| SafeChainBackbone=DeepSeek-R1-Distill-Llama-8B2025.09 | — | 44.5 | |
| SafeChainBackbone=DeepSeek-R1-Distill-Qwen-7B2025.09 | — | 49 | |
| SafeKeyBackbone=DeepSeek-R1-Distill-Llama-8B2025.09 | — | 42.9 | |
| SafeKeyBackbone=DeepSeek-R1-Distill-Qwen-7B2025.09 | — | 51.5 | |
| Segment Selective SFTModel Backbone=Qwen2.5-7B-Instruct, Decoding Strategy=Temperature Sampling2026.01 | — | 43.7 | |
| STARBackbone=DeepSeek-R1-Distill-Llama-8B2025.09 | — | 47 | |
| STARBackbone=DeepSeek-R1-Distill-Qwen-7B2025.09 | — | 49 |