Knowledge Evaluation on GPQA
90.4AccuracyCoT2-Meta
Evaluation Results
| Method | Links | |
|---|---|---|
| CoT2-MetaStrategy=Ours (CoT2-Meta), Inference Budget=C=162026.03 | 90.4 | |
| ReST-MCTS*Strategy=ReST-MCTS*, Inference Budget=C=162026.03 | 85.2 | |
| Vanilla ToTStrategy=Vanilla ToT, Inference Budget=C=162026.03 | 83.5 | |
| Best-of-16Strategy=Best-of-16, Inference Budget=C=162026.03 | 79.6 | |
| Greedy CoTStrategy=Greedy CoT, Inference Budget=C=162026.03 | 74.2 | |
| DeepSeek-R1-Distill-Qwen-32B (Reasoning)Model Family=Qwen2.5-32B2026.01 | 59.39 | |
| Llama-3.3-70B-InstructParameters=70B2026.01 | 56.5 | |
| ReasonAnyModel Family=Qwen2.5-32B2026.01 | 56.25 | |
| DictaLM 3.0 24B-ThinkParameters=24B, Variant=Thinking2026.02 | 55.13 | |
| Phi-42026.01 | 50.9 | |
| ReasoningBase Model=Llama-3.1-8B2026.01 | 47.51 | |
| LEDModel Family=Qwen2.5-32B2026.01 | 46.97 | |
| Gemma 3 27BParameters=27B2026.02 | 45 | |
| Mistral Small 3.12026.02 | 44.87 | |
| Task ArithmeticMerging Protocol=Task Arithmetic2026.01 | 42.42 | |
| Qwen2.5-32B-Instruct (Safety)Model Family=Qwen2.5-32B2026.01 | 41.67 | |
| ReasonAnyMerging Protocol=ReasonAny2026.01 | 41.06 | |
| LinearMerging Protocol=Linear2026.01 | 40.1 | |
| LEDMerging Protocol=LED2026.01 | 39.38 | |
| LinearModel Family=Qwen2.5-32B2026.01 | 38.64 | |
| DAREModel Family=Qwen2.5-32B2026.01 | 38.64 | |
| Task ArithmeticModel Family=Qwen2.5-32B2026.01 | 36.36 | |
| DAREMerging Protocol=DARE2026.01 | 36.36 | |
| Qwen3-4BTotal Parameters=4B, Active Parameters=4B, Trained Tokens=36T2025.11 | 34.85 | |
| Multi ModelModel=Qwen3-8B-Base, KV Sharing=X2026.02 | 34.3 | |
| TIESModel Family=Qwen2.5-32B2026.01 | 34.09 | |
| ICaRusModel=Qwen3-8B-Base, KV Sharing=O2026.02 | 33.8 | |
| FuseLLMModel Family=Qwen2.5-32B2026.01 | 33.33 | |
| TIESMerging Protocol=TIES2026.01 | 33.33 | |
| Foundation-Sec-8B-InstructParameters=8B2026.01 | 31.9 | |
| Foundation-Sec-8B-ReasoningParameters=8B2026.01 | 31.7 | |
| Llama-3.2-3BTotal Parameters=3.2B, Active Parameters=3.2B, Trained Tokens=9T2025.11 | 30.6 | |
| Gemma-3-4BTotal Parameters=4B, Active Parameters=4B, Trained Tokens=4T2025.11 | 29.51 | |
| LFM2-8B-A1BTotal Parameters=8.3B, Active Parameters=1.5B, Trained Tokens=13T2025.11 | 29.29 | |
| Llama-Primus-Merged2026.01 | 29.2 | |
| ICaRusModel=LLaMA-3.1-8B, KV Sharing=O2026.02 | 28.8 | |
| ICaRus# Model=3, KV Sharing=O2026.02 | 28.8 | |
| Multi ModelModel=LLaMA-3.1-8B, KV Sharing=X2026.02 | 27.3 | |
| IF Model# Model=1, KV Sharing=.2026.02 | 27.2 | |
| Multi Model# Model=3, KV Sharing=X2026.02 | 27.2 | |
| LFM2-2.6BTotal Parameters=2.6B, Active Parameters=2.6B, Trained Tokens=11T2025.11 | 26.57 | |
| Granite-4.0-HTotal Parameters=7B, Active Parameters=1B, Trained Tokens=15T2025.11 | 26.46 | |
| SmolLM3-3BTotal Parameters=3.1B, Active Parameters=3.1B, Trained Tokens=11T2025.11 | 26.31 | |
| Llama-3.1-8B-InstructParameters=8B2026.01 | 25.7 | |
| FuseLLMMerging Protocol=FuseLLM2026.01 | 24.24 | |
| Base ModelModel=Qwen3-8B-Base, KV Sharing=.2026.02 | 24.2 | |
| SafetyBase Model=Llama-3.1-8B2026.01 | 23.48 | |
| Coding Model# Model=1, KV Sharing=.2026.02 | 21.7 | |
| Math Model# Model=1, KV Sharing=.2026.02 | 20.7 | |
| Base ModelModel=LLaMA-3.1-8B, KV Sharing=.2026.02 | 16.7 | |
| Base Model# Model=1, KV Sharing=.2026.02 | 16.7 |