Question Answering on GPQA (test)
70.4AccuracyTeacher
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| Teacher2026.05 | 70.4 | — | — | — | — | — | |
| UPAVenue=-, Execution Engine=GPT-4o-mini2026.01 | 45.5 | — | — | — | — | — | |
| kNN-MoEBackbone=GPT-OSS2026.01 | 45.45 | — | — | — | — | — | |
| kNN-MoEBackbone=Qwen32026.01 | 44.95 | — | — | — | — | — | |
| 5-shotBackbone=Qwen32026.01 | 44.44 | — | — | — | — | — | |
| Zero-shotBackbone=GPT-OSS2026.01 | 43.94 | — | — | — | — | — | |
| SFT (Router Only)Backbone=Qwen32026.01 | 43.94 | — | — | — | — | — | |
| OPROVenue=ICLR 24, Execution Engine=GPT-4o-mini2026.01 | 43.3 | — | — | — | — | — | |
| Step-BackVenue=ICLR 24, Execution Engine=GPT-4o-mini2026.01 | 42.4 | — | — | — | — | — | |
| Ours (trajectory-specific release rule)Release Rule=trajectory-specific2026.05 | 42.4 | — | — | — | — | — | |
| Qwen3 8BContext=327682026.02 | 41.96 | — | — | — | — | — | |
| SFTBackbone=GPT-OSS2026.01 | 41.92 | — | — | — | — | — | |
| SPOVenue=EMNLP 26, Execution Engine=GPT-4o-mini2026.01 | 41.8 | — | — | — | — | — | |
| CoTVenue=NeurIPS 22, Execution Engine=GPT-4o-mini2026.01 | 41.6 | — | — | — | — | — | |
| SFT (Router Only)Backbone=GPT-OSS2026.01 | 41.41 | — | — | — | — | — | |
| Zero-shotBackbone=Qwen32026.01 | 41.41 | — | — | — | — | — | |
| PromptAgentVenue=ICLR 24, Execution Engine=GPT-4o-mini2026.01 | 41.3 | — | — | — | — | — | |
| Fixed-prefix OPDPrefix Length=20482026.05 | 41.2 | — | — | — | — | — | |
| APEVenue=ICLR 23, Execution Engine=GPT-4o-mini2026.01 | 41.1 | — | — | — | — | — | |
| PromptBreederVenue=ICML 24, Execution Engine=GPT-4o-mini2026.01 | 40.9 | — | — | — | — | — | |
| OPD2026.05 | 40.7 | — | — | — | — | — | |
| Fixed-prefix OPDPrefix Length=10242026.05 | 40.7 | — | — | — | — | — | |
| QWEN3-14BModel Architecture=QWEN3-14B, Configuration Type=Dense, Evaluation Protocol=5-shot, Sparsity Level=0%2026.01 | 40.34 | — | — | — | — | — | |
| RaRVenue=arXiv 23, Execution Engine=GPT-4o-mini2026.01 | 40.2 | — | — | — | — | — | |
| TextGradVenue=Nature 25, Execution Engine=GPT-4o-mini2026.01 | 40.2 | — | — | — | — | — | |
| SFTBackbone=Qwen32026.01 | 39.9 | — | — | — | — | — | |
| IOVenue=-, Execution Engine=GPT-4o-mini2026.01 | 38.9 | — | — | — | — | — | |
| Fixed-prefix OPDPrefix Length=40962026.05 | 38.6 | — | — | — | — | — | |
| Fixed-prefix OPDPrefix Length=81922026.05 | 38.6 | — | — | — | — | — | |
| Llama3.1 7BContext=80962026.02 | 37.71 | — | — | — | — | — | |
| Random release2026.05 | 37.4 | — | — | — | — | — | |
| MCDIFFUSEContext=10242026.02 | 37.19 | — | — | — | — | — | |
| 5-shotBackbone=GPT-OSS2026.01 | 36.87 | — | — | — | — | — | |
| QWEN3-4BModel Architecture=QWEN3-4B, Configuration Type=Dense, Evaluation Protocol=5-shot, Sparsity Level=0%2026.01 | 36.64 | — | — | — | — | — | |
| ExOPD2026.05 | 36.6 | — | — | — | — | — | |
| ICL_bioselection_strategy=ICL, example_type=bio2024.04 | 36.2 | — | — | — | — | — | |
| Qwen2.5 7BContext=327682026.02 | 35.94 | — | — | — | — | — | |
| ICL_sameselection_strategy=ICL, example_type=same2024.04 | 35.8 | — | — | — | — | — | |
| DEEPSEEK-R1-QWEN-8BModel Architecture=DEEPSEEK-R1-QWEN-8B, Configuration Type=Dense, Evaluation Protocol=5-shot, Sparsity Level=0%2026.01 | 35.18 | — | — | — | — | — | |
| N/Aselection_strategy=none2024.04 | 34.4 | — | — | — | — | — | |
| Random_sameselection_strategy=random, example_type=same2024.04 | 33.7 | — | — | — | — | — | |
| Random_diffselection_strategy=random, example_type=different2024.04 | 33.1 | — | — | — | — | — | |
| Random_bioselection_strategy=random, example_type=bio2024.04 | 32.6 | — | — | — | — | — | |
| base2025.01 | 32.11 | — | — | — | — | — | |
| GPT-42025.01 | 32.1 | — | — | — | — | — | |
| TAMA2025.01 | 31.92 | — | — | — | — | — | |
| MISTRAL-7BModel Architecture=MISTRAL-7B, Configuration Type=Dense, Evaluation Protocol=5-shot, Sparsity Level=0%2026.01 | 31.8 | — | — | — | — | — | |
| Relevantselection_strategy=relevant examples2024.04 | 31.6 | — | — | — | — | — | |
| LLAMA-3.1-8BModel Architecture=LLAMA-3.1-8B, Configuration Type=Dense, Evaluation Protocol=5-shot, Sparsity Level=0%2026.01 | 31.38 | — | — | — | — | — | |
| Student2026.05 | 31.1 | — | — | — | — | — | |
| DARTModel Architecture=QWEN3-4B, Configuration Type=Sparse, Evaluation Protocol=5-shot, Sparsity Level=70%2026.01 | 30.97 | — | — | — | — | — | |
| DARTModel Architecture=QWEN3-14B, Configuration Type=Sparse, Evaluation Protocol=5-shot, Sparsity Level=70%2026.01 | 30.59 | — | — | — | — | — | |
| GPT-3.52025.01 | 29.8 | — | — | — | — | — | |
| kNN-MoEBackbone=OLMoE2026.01 | 29.8 | — | — | — | — | — | |
| LLAMA-3.2-3BModel Architecture=LLAMA-3.2-3B, Configuration Type=Dense, Evaluation Protocol=5-shot, Sparsity Level=0%2026.01 | 29.63 | — | — | — | — | — | |
| Zero-shotBackbone=OLMoE2026.01 | 27.27 | — | — | — | — | — | |
| DARTModel Architecture=DEEPSEEK-R1-QWEN-8B, Configuration Type=Sparse, Evaluation Protocol=5-shot, Sparsity Level=70%2026.01 | 27 | — | — | — | — | — | |
| DARTModel Architecture=MISTRAL-7B, Configuration Type=Sparse, Evaluation Protocol=5-shot, Sparsity Level=70%2026.01 | 26.16 | — | — | — | — | — | |
| DARTModel Architecture=LLAMA-3.1-8B, Configuration Type=Sparse, Evaluation Protocol=5-shot, Sparsity Level=70%2026.01 | 25.9 | — | — | — | — | — | |
| DEEPSEEK-R1-LLAMA-8BModel Architecture=DEEPSEEK-R1-LLAMA-8B, Configuration Type=Dense, Evaluation Protocol=5-shot, Sparsity Level=0%2026.01 | 25.57 | — | — | — | — | — | |
| SFT (Router Only)Backbone=OLMoE2026.01 | 24.24 | — | — | — | — | — | |
| DARTModel Architecture=LLAMA-3.2-3B, Configuration Type=Sparse, Evaluation Protocol=5-shot, Sparsity Level=70%2026.01 | 24.13 | — | — | — | — | — | |
| DARTModel Architecture=DEEPSEEK-R1-LLAMA-8B, Configuration Type=Sparse, Evaluation Protocol=5-shot, Sparsity Level=70%2026.01 | 22.77 | — | — | — | — | — | |
| 5-shotBackbone=OLMoE2026.01 | 21.72 | — | — | — | — | — | |
| SFTBackbone=OLMoE2026.01 | 21.72 | — | — | — | — | — | |
| Agentic Reasoning w/DeepSeek-R1Reasoning Strategy=Agentic Reasoning, Backbone=DeepSeek-R12025.02 | — | 94.5 | 73.7 | 80.5 | 81.2 | — | |
| Agentic Reasoning w/QwQ-32BReasoning Strategy=Agentic Reasoning, Backbone=QwQ-32B2025.02 | — | 88.1 | 58.3 | 79.6 | 69.7 | — | |
| BaseBase Model=DeepSeek-Math-7B-RL, Sampling Strategy=Standard Decoding2026.01 | — | — | — | — | 33.3 | — | |
| Best-of-NBase Model=DeepSeek-Math-7B-RL, Sampling Strategy=Best-of-N2026.01 | — | — | — | — | 29.7 | — | |
| DeepSeek-R1Reasoning Strategy=Direct Reasoning, Backbone=DeepSeek-R12025.02 | — | 86.8 | 56.1 | 63.8 | 71.5 | — | |
| DeepSeek-R1-LiteReasoning Strategy=Direct Reasoning2025.01 | — | — | — | — | 58.5 | — | |
| GPT-4oReasoning Strategy=Direct Reasoning, Backbone=GPT-4o2025.02 | — | 59.5 | 40.2 | 61.6 | 50 | — | |
| GPT-4oReasoning Strategy=Direct Reasoning2025.01 | — | 59.5 | 40.2 | 61.6 | 50.6 | — | |
| Llama3.3-70BReasoning Strategy=Direct Reasoning, Backbone=Llama3.3-70B2025.02 | — | 54.7 | 31.2 | 52.6 | 43.4 | — | |
| Llama3.3-70BReasoning Strategy=Direct Reasoning, Backbone=Llama3.3-70B-Instruct2025.01 | — | 54.7 | 31.2 | 52.6 | 43.4 | — | |
| Low-temperatureBase Model=DeepSeek-Math-7B-RL, Sampling Strategy=Low-temperature2026.01 | — | — | — | — | 30.3 | — | |
| MCMC Power SamplingBase Model=DeepSeek-Math-7B-RL, Sampling Strategy=MCMC Power Sampling2026.01 | — | — | — | — | 34.9 | — | |
| o1Reasoning Strategy=Direct Reasoning, Backbone=o12025.02 | — | 92.8 | 64.7 | 69.2 | 78 | — | |
| o1-previewReasoning Strategy=Direct Reasoning2025.01 | — | 89.4 | 59.9 | 65.9 | 73.3 | — | |
| o3-mini-highReasoning Strategy=Direct Reasoning, Backbone=o3-mini, Configuration=high2025.02 | — | — | — | — | 79.7 | — | |
| o3-mini-lowReasoning Strategy=Direct Reasoning, Backbone=o3-mini, Configuration=low2025.02 | — | — | — | — | 70.6 | — | |
| o3-mini-midReasoning Strategy=Direct Reasoning, Backbone=o3-mini, Configuration=mid2025.02 | — | — | — | — | 76.8 | — | |
| Power SamplingBase Model=DeepSeek-Math-7B-RL, Sampling Strategy=Power Sampling (ours)2026.01 | — | — | — | — | 36.4 | — | |
| Qwen2.5-32BReasoning Strategy=Direct Reasoning, Backbone=Qwen2.5-32B-Instruct2025.01 | — | 57 | 33.3 | 52.6 | 45.5 | — | |
| Qwen2.5-72BReasoning Strategy=Direct Reasoning, Backbone=Qwen2.5-72B-Instruct2025.01 | — | 57 | 37.6 | 68.4 | 49 | — | |
| Qwen2.5-Coder-32BReasoning Strategy=Direct Reasoning, Backbone=Qwen2.5-Coder-32B-Instruct2025.01 | — | 37.2 | 25.8 | 57.9 | 33.8 | — | |
| QwQ-32BReasoning Strategy=Direct Reasoning, Backbone=QwQ-32B2025.02 | — | 75.6 | 39.8 | 68.4 | 58.1 | — | |
| QwQ-32BReasoning Strategy=Direct Reasoning, Backbone=QwQ-32B-Preview2025.01 | — | 75.6 | 39.8 | 68.4 | 58.1 | — | |
| QwQ-32BStrategy=Direct2026.02 | — | — | — | — | — | 43.4 | |
| RAG-Qwen2.5-32BReasoning Strategy=Standard RAG, Backbone=Qwen2.5-32B-Instruct2025.01 | — | 57 | 37.6 | 52.6 | 47.5 | — | |
| RAG-QwQ-32BReasoning Strategy=Standard RAG, Backbone=QwQ-32B-Preview2025.01 | — | 76.7 | 38.7 | 73.7 | 58.6 | — | |
| RAG-QwQ-32BStrategy=RAG2026.02 | — | — | — | — | — | 64.6 | |
| RAgent w/DeepSeek-R1Reasoning Strategy=Retrieve/Search in Reasoning, Backbone=DeepSeek-R1, Framework=RAgent2025.02 | — | 87.7 | 58.2 | 65.7 | 72.9 | — | |
| RAgent w/QwQ-32BReasoning Strategy=Retrieve/Search in Reasoning, Backbone=QwQ-32B, Framework=RAgent2025.02 | — | 76.7 | 46.2 | 68.4 | 61.6 | — | |
| RAgent-Qwen2.5-32BReasoning Strategy=RAG Agent, Backbone=Qwen2.5-32B-Instruct2025.01 | — | 58.1 | 33.3 | 63.2 | 47 | — | |
| RAgent-QwQ-32BReasoning Strategy=RAG Agent, Backbone=QwQ-32B-Preview2025.01 | — | 76.7 | 46.2 | 68.4 | 61.6 | — | |
| Search-o1Reasoning Strategy=Reason-in-Documents, Backbone=QwQ-32B-Preview2025.01 | — | 77.9 | 47.3 | 78.9 | 63.6 | — | |
| Search-o1-32BBackbone=o1-32B2026.02 | — | — | — | — | — | 67.2 | |
| SearchO1 w/DeepSeek-R1Reasoning Strategy=Retrieve/Search in Reasoning, Backbone=DeepSeek-R1, Framework=SearchO12025.02 | — | 90.2 | 61.3 | 71.4 | 74.6 | — | |
| SearchO1 w/QwQ-32BReasoning Strategy=Retrieve/Search in Reasoning, Backbone=QwQ-32B, Framework=SearchO12025.02 | — | 77.9 | 47.3 | 78.9 | 63.6 | — |