Science Question Answering on GPQA Diamond
91.9AccuracyGemini-3.0
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Gemini-3.0variant=Pro, protocol=Pass@12025.12 | 91.9 | 8,000 | |
| Gemma-4 31BEvaluation Mode=Reasoning+Rotate2Think2026.06 | 86.87 | 5,614 | |
| GPT-5variant=High, protocol=Pass@12025.12 | 85.7 | 8,000 | |
| DeepSeek-V3.2variant=Speciale, protocol=Pass@12025.12 | 85.7 | 16,000 | |
| Kimi-K2variant=Thinking, protocol=Pass@12025.12 | 84.5 | 12,000 | |
| Gemma-4 31BEvaluation Mode=Reasoning2026.06 | 84.34 | 5,728 | |
| DeepSeek-V3.2variant=Thinking, protocol=Pass@12025.12 | 82.4 | 7,000 | |
| Gemma-4 31BEvaluation Mode=Base+Rotate2Think2026.06 | 78.28 | 1,390 | |
| Gemma-4 31BEvaluation Mode=Base2026.06 | 74.24 | 1,284 | |
| ϕ-Decoding + SCBackbone=DeepSeek-R1-0528-Qwen3-8B2026.05 | 72.6 | 95.8 | |
| Phi-4 14BEvaluation Mode=Reasoning+Rotate2Think2026.06 | 69.7 | 11,569 | |
| DDCBackbone=DeepSeek-R1-0528-Qwen3-8B2026.05 | 69.6 | 13.9 | |
| Predictive Decoding + SCBackbone=DeepSeek-R1-0528-Qwen3-8B2026.05 | 69.5 | 98.6 | |
| Qwen3 4BEvaluation Mode=Reasoning+Rotate2Think2026.06 | 69.19 | 9,126 | |
| Phi-4 14BEvaluation Mode=Reasoning2026.06 | 69.19 | 11,353 | |
| SHAPEBackbone=Qwen3-30B-A3B, Pruning Ratio=40%2026.06 | 67.68 | — | |
| GPT-4.12025.07 | 67.51 | — | |
| Qwen3 4BEvaluation Mode=Reasoning2026.06 | 67.17 | 8,018 | |
| LessIsMoreModel=Qwen3-14B, Token Budget=6K, Sampled answers=162025.08 | 65.15 | — | |
| LessIsMoreModel=Qwen3-14B, Token Budget=4K, Sampled answers=162025.08 | 64.85 | — | |
| SMCS2025.07 | 64.81 | — | |
| Self-MoA2025.07 | 64.65 | — | |
| Full AttnModel=Qwen3-14B, Sampled answers=162025.08 | 64.02 | — | |
| Gemma-4 E4BEvaluation Mode=Base+Rotate2Think2026.06 | 64 | 3,254 | |
| LessIsMoreModel=Qwen3-14B, Token Budget=2K, Sampled answers=162025.08 | 63.61 | — | |
| LessIsMoreModel=Qwen3-14B, Token Budget=1K, Sampled answers=162025.08 | 63.42 | — | |
| SHAPEBackbone=GPT-OSS-20B, Pruning Ratio=20%2026.06 | 62.63 | — | |
| GPT-OSS-20BSetting=Baseline2026.06 | 61.62 | — | |
| LessIsMoreModel=Qwen3-8B, Token Budget=4K, Sampled answers=162025.08 | 61.58 | — | |
| LessIsMoreModel=Qwen3-8B, Token Budget=6K, Sampled answers=162025.08 | 61.11 | — | |
| LessIsMoreModel=Qwen3-8B, Token Budget=2K, Sampled answers=162025.08 | 60.65 | — | |
| Full AttnModel=Qwen3-8B, Sampled answers=162025.08 | 60.54 | — | |
| Gemma-4 E4BEvaluation Mode=Reasoning+Rotate2Think2026.06 | 60.33 | 3,342 | |
| SHAPEBackbone=Qwen3-30B-A3B, Pruning Ratio=20%2026.06 | 60.1 | — | |
| DDCBackbone=Qwen3-4B2026.05 | 59.1 | 13.7 | |
| Gemma-4 E4BEvaluation Mode=Reasoning2026.06 | 58.67 | 3,646 | |
| LessIsMoreModel=Qwen3-8B, Token Budget=1K, Sampled answers=162025.08 | 58.62 | — | |
| Qwen3-30B-A3BSetting=Baseline2026.06 | 58.59 | — | |
| QwQ-32B2025.07 | 57.24 | — | |
| LessIsMoreModel=Qwen3-4B, Token Budget=4K, Sampled answers=162025.08 | 56.84 | — | |
| LessIsMoreModel=Qwen3-4B, Token Budget=6K, Sampled answers=162025.08 | 56.64 | — | |
| ϕ-Decoding + SCBackbone=Qwen3-4B2026.05 | 56.6 | 86.9 | |
| Phi-4 14BEvaluation Mode=Base+Rotate2Think2026.06 | 56.57 | 746 | |
| SHAPEBackbone=GPT-OSS-20B, Pruning Ratio=40%2026.06 | 56.57 | — | |
| LessIsMoreModel=Qwen3-4B, Token Budget=1K, Sampled answers=162025.08 | 56.48 | — | |
| LessIsMoreModel=Qwen3-4B, Token Budget=2K, Sampled answers=162025.08 | 56.23 | — | |
| Full AttnModel=Qwen3-4B, Sampled answers=162025.08 | 56.19 | — | |
| Gemma-4 E4BEvaluation Mode=Base2026.06 | 55.07 | 1,979 | |
| Predictive Decoding + SCBackbone=Qwen3-4B2026.05 | 54.8 | 94.3 | |
| Phi-4 14BEvaluation Mode=Base2026.06 | 54.55 | 658 | |
| Qwen3 4BEvaluation Mode=Base+Rotate2Think2026.06 | 48.48 | 549 | |
| Qwen3 4BEvaluation Mode=Base2026.06 | 46.67 | 737 | |
| LessIsMoreModel=DeepSeek-8B, Token Budget=6K, Sampled answers=162025.08 | 43.31 | — | |
| LessIsMoreModel=DeepSeek-8B, Token Budget=4K, Sampled answers=162025.08 | 43.08 | — | |
| Full AttnModel=DeepSeek-8B, Sampled answers=162025.08 | 42.8 | — | |
| GRPOBase model=Qwen2.5-Math-7B2025.07 | 40.4 | — | |
| RL-PLUSBase model=Qwen2.5-Math-7B2025.07 | 40.4 | — | |
| LessIsMoreModel=DeepSeek-8B, Token Budget=2K, Sampled answers=162025.08 | 39.11 | — | |
| LessIsMoreModel=DeepSeek-8B, Token Budget=1K, Sampled answers=162025.08 | 35.74 | — | |
| PLM-HoneyBee-8BModel scale=7B-8B2025.10 | 33.3 | — | |
| SHAPEBackbone=DeepSeek-V2-Lite, Pruning Ratio=20%2026.06 | 31.52 | — | |
| DeepSeek-V2-LiteSetting=Baseline2026.06 | 28.28 | — | |
| PLM-HoneyBee-3BModel scale=3B-4B2025.10 | 27.7 | — | |
| SHAPEBackbone=DeepSeek-V2-Lite, Pruning Ratio=40%2026.06 | 26.53 | — | |
| InternVL-2.5-8BModel scale=7B-8B2025.10 | 26.3 | — | |
| Qwen2.5-VL-7B-InstructModel scale=7B-8B2025.10 | 26.3 | — | |
| PLM-HoneyBee-1BModel scale=1B2025.10 | 25.8 | — | |
| InternVL-2.5-4BModel scale=3B-4B2025.10 | 25.3 | — | |
| SFTBase model=Qwen2.5-Math-7B2025.07 | 24.7 | — | |
| SFT+GRPOBase model=Qwen2.5-Math-7B2025.07 | 24.2 | — | |
| Repeated SamplingModel=Llama-3.2-3B-Instruct, Decoding=Majority Voting2025.10 | 23.23 | — | |
| GUIDEDSAMPLINGModel=Llama-3.2-3B-Instruct, Decoding=Majority Voting2025.10 | 23.23 | — | |
| Qwen2.5-VL-3B-InstructModel scale=3B-4B2025.10 | 22.7 | — | |
| Repeated SamplingModel=Qwen2.5-3B-Instruct, Decoding=Majority Voting2025.10 | 20.71 | — | |
| GUIDEDSAMPLINGModel=Qwen2.5-3B-Instruct, Decoding=Majority Voting2025.10 | 20.2 | — | |
| InternVL-3-8B-InstructModel scale=7B-8B2025.10 | 20.2 | — | |
| InternVL-3-1B-InstructModel scale=1B2025.10 | 19.2 | — | |
| Tree-of-thoughtModel=Llama-3.2-3B-Instruct, Decoding=Majority Voting2025.10 | 19.19 | — | |
| PLM-3BModel scale=3B-4B2025.10 | 18.7 | — | |
| PLM-8BModel scale=7B-8B2025.10 | 16.2 | — | |
| Base ModelBase model=Qwen2.5-Math-7B2025.07 | 13.1 | — | |
| InternVL-2.5-1BModel scale=1B2025.10 | 12.1 | — | |
| PLM-1BModel scale=1B2025.10 | 7.1 | — | |
| Tree-of-thoughtModel=Qwen2.5-3B-Instruct, Decoding=Majority Voting2025.10 | 7.07 | — |