Commonsense Reasoning on StrategyQA
95.7AccuracyREBALANCE
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| REBALANCEBackbone=QwQ-32B2026.03 | 95.7 | 265 | |
| REBALANCEBackbone=Qwen3-14B2026.03 | 94.3 | 260 | |
| BaselineBackbone=Qwen3-14B2026.03 | 94.2 | 267 | |
| BaselineBackbone=QwQ-32B2026.03 | 93.6 | 274 | |
| ORACLEBackbone=Qwen-2.5-7B-instruct, Setting=zero-shot2026.03 | 92.6 | — | |
| Claude 3.5 SonnetEvaluation Protocol=Closed-Source API, Emission (gCO2/q)=3.85†, Throughput (Tok/s)=62.12026.03 | 92.2 | — | |
| GPT-4oEvaluation Protocol=Closed-Source API, Emission (gCO2/q)=4.52†, Throughput (Tok/s)=55.22026.03 | 92 | — | |
| ORACLEBackbone=Llama-3.1-8B-Instruct, Setting=zero-shot2026.03 | 91.8 | — | |
| self-rewardingBackbone=Qwen-2.5-7B-instruct, Setting=zero-shot2026.03 | 91.5 | — | |
| Gemini 2.5 ProEvaluation Protocol=Closed-Source API, Emission (gCO2/q)=3.90†, Throughput (Tok/s)=58.42026.03 | 91.5 | — | |
| self-rewardingBackbone=Llama-3.1-8B-Instruct, Setting=zero-shot2026.03 | 90.8 | — | |
| EcoThinkEvaluation Protocol=Adaptive Inference, Emission (gCO2/q)=1.32, Throughput (Tok/s)=148.62026.03 | 90.5 | — | |
| PaLM 2Number of exemplars (k-shot)=6, Model variant=Instruction-tuned, Chain-of-thought prompting (CoT)=true, Self-consistency (SC)=true2023.05 | 90.4 | — | |
| CoTBackbone=Llama-3.1-8B-Instruct, Setting=zero-shot2026.03 | 89.9 | — | |
| RFTBackbone=Llama-3.1-8B-Instruct, Setting=zero-shot2026.03 | 89.9 | — | |
| ORACLEBackbone=Mistral-7B-Instruct-v0.3, Setting=zero-shot2026.03 | 89.9 | — | |
| CoTBackbone=Qwen-2.5-7B-instruct, Setting=four-shot2026.03 | 89.9 | — | |
| Clean ModelTarget Model=Mistral-7B, Trigger Condition=Without Trigger2025.04 | 89.4 | — | |
| ShadowCoTTarget Model=Mistral-7B, Trigger Condition=Without Trigger2025.04 | 89.1 | — | |
| RFTBackbone=Mistral-7B-Instruct-v0.3, Setting=zero-shot2026.03 | 89 | — | |
| self-rewardingBackbone=Mistral-7B-Instruct-v0.3, Setting=zero-shot2026.03 | 89 | — | |
| REBALANCEBackbone=DeepSeek-R1-Distill-Qwen-7B2026.03 | 88.9 | 310 | |
| CoTBackbone=Llama-3.1-8B-Instruct, Setting=four-shot2026.03 | 88.6 | — | |
| ToT-SFTBackbone=Qwen-2.5-7B-instruct, Setting=zero-shot2026.03 | 88.2 | — | |
| BaselineBackbone=DeepSeek-R1-Distill-Qwen-7B2026.03 | 88.1 | 350 | |
| CoTBackbone=Mistral-7B-Instruct-v0.3, Setting=four-shot2026.03 | 88.1 | — | |
| DarkMindTarget Model=Mistral-7B, Trigger Condition=Without Trigger2025.04 | 88 | — | |
| CoTBackbone=Mistral-7B-Instruct-v0.3, Setting=zero-shot2026.03 | 87.3 | — | |
| RFTBackbone=Qwen-2.5-7B-instruct, Setting=zero-shot2026.03 | 87.1 | — | |
| BadChainTarget Model=Mistral-7B, Trigger Condition=Without Trigger2025.04 | 87.1 | — | |
| CoTBackbone=Qwen-2.5-7B-instruct, Setting=zero-shot2026.03 | 86.9 | — | |
| FrugalGPT (Cascade)Evaluation Protocol=Standard CoT, Emission (gCO2/q)=1.95, Throughput (Tok/s)=88.52026.03 | 86.9 | — | |
| ToT-SFTBackbone=Llama-3.1-8B-Instruct, Setting=zero-shot2026.03 | 85.2 | — | |
| Qwen-3-8B-InstructEvaluation Protocol=Standard CoT, Emission (gCO2/q)=2.12, Throughput (Tok/s)=98.52026.03 | 84.8 | — | |
| ToT-SFTBackbone=Mistral-7B-Instruct-v0.3, Setting=zero-shot2026.03 | 84.4 | — | |
| ToT (GPT-4)Base Model=GPT-42024.03 | 83 | — | |
| PaLM + CoT + SCNumber of exemplars (k-shot)=6, Chain-of-thought prompting (CoT)=true, Self-consistency (SC)=true2023.05 | 81.6 | — | |
| Self-consistencyModel=PaLM-540B2022.03 | 81.6 | — | |
| Llama-3.1-8B-InstructEvaluation Protocol=Standard CoT, Emission (gCO2/q)=2.15, Throughput (Tok/s)=95.82026.03 | 81.4 | — | |
| Self-consistencyModel=GPT-3 (Code-davinci-002)2022.03 | 79.8 | — | |
| Self-consistency (Code-davinci-002)Base Model=Code-davinci-0022024.03 | 79.8 | — | |
| Few-shot-CoT (GPT-4)Base Model=GPT-4, Prompting Strategy=Few-shot-CoT2024.03 | 79.1 | — | |
| Few-shot-CoT-CP (GPT-4) + SCBase Model=GPT-4, Prompting Strategy=Few-shot-CoT-CP, Self-Consistency=true2024.03 | 78.8 | — | |
| DIVERSEModel=code-davinci-0022022.06 | 78.6 | — | |
| Few-shot-CoT-CP (GPT-4)Base Model=GPT-4, Prompting Strategy=Few-shot-CoT-CP2024.03 | 78.2 | — | |
| TADBModel=GPT-4 Turbo2024.07 | 78 | — | |
| Self-ConsistencyModel=code-davinci-0022022.06 | 77.6 | — | |
| QAP25Model=GPT-4 Turbo, Constraint (N)=252024.07 | 77.6 | — | |
| QAP150Model=GPT-4 Turbo, Constraint (N)=1502024.07 | 77.6 | — | |
| QAP100Model=GPT-4 Turbo, Constraint (N)=1002024.07 | 77.2 | — | |
| PS+Model=GPT-4 Turbo2024.07 | 77.1 | — | |
| QAP50Model=GPT-4 Turbo, Constraint (N)=502024.07 | 76.9 | — | |
| BaselineModel=GPT-4 Turbo2024.07 | 76.3 | — | |
| SCMoEBackbone=Mixtral 8x7B, Routing strategy=rank-2 routing2024.05 | 76.29 | — | |
| QAP200Model=GPT-4 Turbo, Constraint (N)=2002024.07 | 75.9 | — | |
| Greedy DecodeModel=PaLM 540B2022.06 | 75.3 | — | |
| CoT-promptingModel=PaLM-540B2022.03 | 75.3 | — | |
| CoTModel=GPT-4 Turbo2024.07 | 75.1 | — | |
| Contrastive SearchBackbone=Mixtral 8x7B2024.05 | 74.85 | — | |
| DIVERSEModel=text-davinci-0022022.06 | 74.8 | — | |
| Contrastive DecodingBackbone=Mixtral 8x7B, Amateur Model=Mistral-7B [16]2024.05 | 74.45 | — | |
| Dynamic RoutingBackbone=Mixtral 8x7B2024.05 | 74.41 | — | |
| Ensemble RoutingBackbone=Mixtral 8x7B2024.05 | 74.37 | — | |
| D-RPCStudent Model=Llama 3.1 8B Instruct2026.05 | 74.15 | — | |
| SOTA (Fine-tuned)Mode=Fine-tuned2022.06 | 73.9 | — | |
| Previous SoTA2022.03 | 73.9 | — | |
| Fine-tuning2023.06 | 73.9 | — | |
| CoT-promptingModel=GPT-3 (Code-davinci-002)2022.03 | 73.4 | — | |
| CoTStudent Model=Llama 3.1 8B Instruct2026.05 | 73.32 | — | |
| COMPACTStudent Architecture=Qwen2.5-7B, Distillation Method=COMPACT2026.01 | 72.93 | — | |
| FireAct (Multi-task + CoT)Method Category=Fine-tuning, Base Model=GPT-3.5, Training Data=Multi-task + CoT2023.10 | 72.9 | — | |
| GreedyBackbone=Mixtral 8x7B2024.05 | 72.83 | — | |
| DCoTStudent Model=Llama 3.1 8B Instruct2026.05 | 72.53 | — | |
| FreeformStudent Model=Llama 3.1 8B Instruct2026.05 | 72.27 | — | |
| Greedy DecodeModel=code-davinci-0022022.06 | 72 | — | |
| COMPACTStudent Architecture=Llama3.1-8B, Distillation Method=COMPACT2026.01 | 71.72 | — | |
| SoftCoTN=10, Backbone=Qwen3-8B2025.02 | 71.18 | — | |
| EDITStudent Architecture=Llama3.1-8B, Distillation Method=EDIT2026.01 | 71.15 | — | |
| SBS-KDStudent Architecture=Llama3.1-8B, Distillation Method=SBS-KD2026.01 | 71.05 | — | |
| DoLaBackbone=Mixtral 8x7B2024.05 | 71.04 | — | |
| Zero-Shot CoTN=10, Backbone=Qwen3-8B2025.02 | 70.96 | — | |
| Zero-Shot Assist-CoTN=10, Backbone=Qwen3-8B2025.02 | 70.92 | — | |
| CommitteeStudent Architecture=Llama3.1-8B, Distillation Method=Committee2026.01 | 70.87 | — | |
| Llama3.1-8B-InstructStudent Architecture=Llama3.1-8B, Distillation Method=Zero-shot2026.01 | 70.74 | — | |
| SGFTStudent Model=Llama 3.1 8B Instruct2026.05 | 70.31 | — | |
| SoftCoTN=1, Backbone=Qwen3-8B2025.02 | 70.17 | — | |
| Zero-Shot CoTN=1, Backbone=Qwen3-8B2025.02 | 69.87 | — | |
| Self-ConsistencyModel=text-davinci-0022022.06 | 69.8 | — | |
| Zero-Shot Assist-CoTN=1, Backbone=Qwen3-8B2025.02 | 69.78 | — | |
| SoftCoTBackbone=LLaMA-3.1-8B-Instruct2025.02 | 69.04 | — | |
| Few-shot-CPBase Model=GPT-3.5-Turbo, Prompting Strategy=Few-shot-CP2024.03 | 68.7 | — | |
| Qwen2.5-7B-InstructStudent Architecture=Qwen2.5-7B, Distillation Method=Zero-shot2026.01 | 68.68 | — | |
| DUPLLM=GPT-3.5-turbo, Evaluation Protocol=Zero-shot2024.04 | 68.5 | — | |
| MCC-KDStudent Architecture=Llama3.1-8B, Distillation Method=MCC-KD2026.01 | 67.99 | — | |
| CoK + F2-VBase Model=text-davinci-0022023.06 | 67.9 | — | |
| Zero-shot-CP + SCBase Model=GPT-3.5-Turbo, Prompting Strategy=Zero-shot-CP, Self-Consistency=true2024.03 | 67.9 | — | |
| Self-consistencyModel=LaMDA-137B2022.03 | 67.8 | — | |
| REBALANCEBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.03 | 67.7 | 401 | |
| EDITStudent Architecture=Qwen2.5-7B, Distillation Method=EDIT2026.01 | 67.5 | — | |
| Zero-shot-CPBase Model=GPT-3.5-Turbo, Prompting Strategy=Zero-shot-CP2024.03 | 67.3 | — |