Question Answering on GPQA (Accuracy)
84.2AccuracyUPA
Evaluation Results
| Method | Links | |
|---|---|---|
| UPAExecutor=GPT-52026.01 | 84.2 | |
| CoTExecutor=GPT-52026.01 | 83.1 | |
| SPOExecutor=GPT-52026.01 | 82.5 | |
| IOExecutor=GPT-52026.01 | 79.1 | |
| M2CLBase Model=Llama-70B, Number of LLMs=42026.02 | 78.7 | |
| UPAExecutor=DeepSeek-V3.22026.01 | 78.3 | |
| UPAExecutor=Claude-4.5-Sonnet2026.01 | 75.8 | |
| CoTExecutor=DeepSeek-V3.22026.01 | 75.5 | |
| IOExecutor=Claude-4.5-Sonnet2026.01 | 74.7 | |
| SPOExecutor=Claude-4.5-Sonnet2026.01 | 73.7 | |
| SPOExecutor=DeepSeek-V3.22026.01 | 73.6 | |
| CoTExecutor=Claude-4.5-Sonnet2026.01 | 73.2 | |
| M2CLBase Model=Qwen-72B, Number of LLMs=42026.02 | 73.1 | |
| IOExecutor=DeepSeek-V3.22026.01 | 72.9 | |
| Thought-ICS-AModel=OSS-120B2026.02 | 69 | |
| theory-guided context selection strategyModel=Qwen3-8B, Selection strategy=top-12026.02 | 65.7 | |
| GenICLModel=Qwen3-8B, Selection strategy=top-12026.02 | 65.2 | |
| M2CLBase Model=Qwen-14B, Number of LLMs=42026.02 | 64.6 | |
| ZeroModel=Qwen3-8B, Selection strategy=top-1, contextual_information=none2026.02 | 64.6 | |
| DICLModel=Qwen3-8B, Selection strategy=top-12026.02 | 64.6 | |
| TopicKModel=Qwen3-8B, Selection strategy=top-12026.02 | 64.1 | |
| Thought-ICS-SModel=OSS-120B2026.02 | 64 | |
| BM25Model=Qwen3-8B, Selection strategy=top-12026.02 | 63.6 | |
| Thought-ICS-AModel=LLaMA-70B2026.02 | 63 | |
| Thought-ICS-SModel=LLaMA-70B2026.02 | 62 | |
| Qwen 3 VL 32B InstructParameters=32B2025.12 | 61.4 | |
| BoNBase Model=Llama-70B, Number of LLMs=42026.02 | 59.6 | |
| M2CLBase Model=Llama-14B, Number of LLMs=42026.02 | 58.9 | |
| M2CLBase Model=Qwen-7B, Number of LLMs=42026.02 | 58.4 | |
| GPTSwarmBase Model=Llama-70B, Number of LLMs=42026.02 | 54.8 | |
| Qwen 3 32BThinking=No, Parameters=32B2025.12 | 54.4 | |
| DyLANBase Model=Llama-70B, Number of LLMs=42026.02 | 53.8 | |
| Primitives-based MAS2026.02 | 53.2 | |
| CoTModel=LLaMA-70B2026.02 | 53 | |
| CoVeModel=OSS-120B2026.02 | 52 | |
| Self-RefineModel=LLaMA-70B2026.02 | 51 | |
| DebateBase Model=Llama-70B, Number of LLMs=42026.02 | 50.2 | |
| MacNetBase Model=Llama-70B, Number of LLMs=42026.02 | 49.1 | |
| Thought-ICS-AModel=OSS-20B2026.02 | 49 | |
| Olmo 3.1 32B InstructStage=Final Instruct 3.12025.12 | 48.6 | |
| Olmo 3.1 32B InstructStage=DPO2025.12 | 47.9 | |
| GPTSwarmBase Model=Qwen-72B, Number of LLMs=42026.02 | 46.6 | |
| DyLANBase Model=Llama-14B, Number of LLMs=42026.02 | 46.5 | |
| OLMo3-7B-Think-DPO + GRPO (FLIP)Reward Model=Qwen3-4B FLIP, Training Method=GRPO2026.02 | 46.4 | |
| Thought-ICS-SModel=OSS-20B2026.02 | 46 | |
| CoVeModel=LLaMA-70B2026.02 | 46 | |
| BoNBase Model=Qwen-72B, Number of LLMs=42026.02 | 45.9 | |
| BridgeBackbone=DS-Qwen-7B, Generation Width=82025.10 | 45.77 | |
| DyLANBase Model=Qwen-72B, Number of LLMs=42026.02 | 45.5 | |
| CoTModel=OSS-120B2026.02 | 45 | |
| Self-RefineModel=OSS-120B2026.02 | 45 | |
| Gemma 3 27BParameters=27B2025.12 | 45 | |
| OLMo3-7B-Think-DPOBase Model=OLMo3-7B2026.02 | 44.9 | |
| M2CLBase Model=Llama-7B, Number of LLMs=42026.02 | 44.7 | |
| Qwen 2.5 32BParameters=32B2025.12 | 44.6 | |
| OLMo3-7B-Think-DPO + GRPOReward Model=Qwen3-4B LLM-as-a-Judge, Training Method=GRPO2026.02 | 44.6 | |
| DS-Qwen-7BBackbone=DS-Qwen-7B2025.10 | 43.94 | |
| P-MatchBackbone=DS-Qwen-7B2025.10 | 43.75 | |
| RLVRBackbone=DS-Qwen-7B2025.10 | 43.56 | |
| theory-guided context selection strategyModel=Llama-3.1-8B, Selection strategy=top-12026.02 | 43.4 | |
| OLMo3-7B-ThinkStatus=Final version after RLVR2026.02 | 43.3 | |
| OLMo3-7B-Think-DPO + GRPO (FLIP)Reward Model=Qwen3-1.7B FLIP, Training Method=GRPO2026.02 | 43.1 | |
| Thought-ICS-SModel=Qwen-32B2026.02 | 43 | |
| Thought-ICS-AModel=Qwen-32B2026.02 | 43 | |
| DICLModel=Llama-3.1-8B, Selection strategy=top-12026.02 | 43 | |
| GenICLModel=Llama-3.1-8B, Selection strategy=top-12026.02 | 43 | |
| TopicKModel=Llama-3.1-8B, Selection strategy=top-12026.02 | 42.9 | |
| OLMo3-7B-Think-DPO + GRPOReward Model=Qwen3-1.7B LLM-as-a-Judge, Training Method=GRPO2026.02 | 42.9 | |
| BM25Model=Llama-3.1-8B, Selection strategy=top-12026.02 | 42.4 | |
| Sigmoid capability boundariesFLOPs budget=10^24, Boundary estimation method=no-split 0.98-quantile sigmoid2026.02 | 42.4 | |
| CoTModel=Qwen-32B2026.02 | 42 | |
| Olmo 3.1 32B InstructStage=SFT2025.12 | 41.3 | |
| MacNetBase Model=Qwen-72B, Number of LLMs=42026.02 | 41.2 | |
| MacNetBase Model=Llama-14B, Number of LLMs=42026.02 | 40.9 | |
| AgentVerse2026.02 | 40.2 | |
| Gemma 2 27BParameters=27B2025.12 | 39.9 | |
| BridgeBackbone=DS-Llama-8B, Generation Width=82025.10 | 39.65 | |
| DebateBase Model=Qwen-72B, Number of LLMs=42026.02 | 39.6 | |
| RLVRBackbone=DS-Llama-8B2025.10 | 39.46 | |
| Self-RefineModel=Qwen-32B2026.02 | 39 | |
| CoVeModel=Qwen-32B2026.02 | 39 | |
| P-MatchBackbone=DS-Llama-8B2025.10 | 38.83 | |
| Self-Refine2026.02 | 38.3 | |
| CoVeModel=Qwen-14B2026.02 | 38 | |
| Thought-ICS-SModel=Qwen-14B2026.02 | 38 | |
| Thought-ICS-AModel=Qwen-14B2026.02 | 38 | |
| BoNBase Model=Llama-14B, Number of LLMs=42026.02 | 37.7 | |
| MAS-GPT2026.02 | 37.6 | |
| GPTSwarmBase Model=Qwen-7B, Number of LLMs=42026.02 | 37.2 | |
| Self-Consistency2026.02 | 37.2 | |
| Single2026.02 | 36.7 | |
| GPTSwarm2026.02 | 36.5 | |
| BoNBase Model=Qwen-7B, Number of LLMs=42026.02 | 36.4 | |
| OLMo 2 32BParameters=32B2025.12 | 36.4 | |
| BoKBackbone=Qwen2.5-Math-7B, Temperature (τ)=0.90, β=0.01, λ=0.12026.02 | 36.36 | |
| BoKBackbone=Qwen2.5-Math-7B, Temperature (τ)=0.25, β=0.02, λ=0.22026.02 | 36.36 | |
| DebateBase Model=Llama-14B, Number of LLMs=42026.02 | 36.3 | |
| DyLAN2026.02 | 36 | |
| SLOTModel=Qwen2.5-Math-7B, Test-Time Adaptation Method=SLOT2025.12 | 35.86 | |
| DS-Llama-8BBackbone=DS-Llama-8B2025.10 | 35.8 |