Logical Reasoning on LogiQA (Accuracy)
80.4AccuracyQwen3-8B-thinking
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3-8B-thinkingModel Parameters=8B, Reasoning Mode=Thinking2026.01 | 80.4 | |
| DenserModel=GPT-4o2025.12 | 79.1 | |
| Process SupervisionModel=GPT-4o2025.12 | 78.9 | |
| Reflection-CoTModel=GPT-4o2025.12 | 78.4 | |
| Self-VerificationModel=GPT-4o2025.12 | 77.8 | |
| Tree-of-ThoughtModel=GPT-4o2025.12 | 77.7 | |
| Process SupervisionModel=DeepSeek-V32025.12 | 77.4 | |
| Self-ConsistencyModel=GPT-4o2025.12 | 77.1 | |
| Reflection-CoTModel=DeepSeek-V32025.12 | 76.9 | |
| DenserModel=DeepSeek-V32025.12 | 76.9 | |
| Think-to-ThinkModel=GPT-4o2025.12 | 76.4 | |
| GPT-5.12026.01 | 76.34 | |
| DenserModel=o1-mini2025.12 | 75.8 | |
| Tree-of-ThoughtModel=DeepSeek-V32025.12 | 75.6 | |
| Self-VerificationModel=DeepSeek-V32025.12 | 75.6 | |
| Process SupervisionModel=o1-mini2025.12 | 75.3 | |
| Self-ConsistencyModel=DeepSeek-V32025.12 | 75 | |
| Reflection-CoTModel=o1-mini2025.12 | 74.8 | |
| Chain-of-ThoughtModel=GPT-4o2025.12 | 74.6 | |
| Self-VerificationModel=o1-mini2025.12 | 74.3 | |
| Think-to-ThinkModel=DeepSeek-V32025.12 | 74.3 | |
| Tree-of-ThoughtModel=o1-mini2025.12 | 73.4 | |
| Qwen3-8BModel Parameters=8B2026.01 | 73.2 | |
| Self-ConsistencyModel=o1-mini2025.12 | 72.8 | |
| Chain-of-ThoughtModel=DeepSeek-V32025.12 | 72.5 | |
| Think-to-ThinkModel=o1-mini2025.12 | 71.9 | |
| Chain-of-ThoughtModel=o1-mini2025.12 | 70.2 | |
| SGR-LLaMA-3.3-70BBackbone=LLaMA-3.3-70B2026.01 | 69.91 | |
| AceMADBase Model=Qwen3-235B-A22B-Instruct, Iteration Rounds (T)=2, Number of Agents (N)=52026.03 | 59.69 | |
| AceMADBase Model=Qwen3-235B-A22B-Instruct, Iteration Rounds (T)=3, Number of Agents (N)=52026.03 | 59.69 | |
| Sparse MADBase Model=Qwen3-235B-A22B-Instruct, Number of Agents (N)=52026.03 | 47.81 | |
| Decentralized MADBase Model=Qwen3-235B-A22B-Instruct, Number of Agents (N)=52026.03 | 45.62 | |
| Centralized MADBase Model=Qwen3-235B-A22B-Instruct, Number of Agents (N)=52026.03 | 45.62 | |
| Majority VotingBase Model=Qwen3-235B-A22B-Instruct, Number of Agents (N)=52026.03 | 43.12 | |
| Single AgentBase Model=Qwen3-235B-A22B-Instruct, Number of Agents (N)=52026.03 | 42.5 | |
| MMKEBackbone=Qwen2.5-7B Instruct, Consolidation=false2025.12 | 41 | |
| MMKE+Con.Backbone=Qwen2.5-7B Instruct, Consolidation=true2025.12 | 39.2 | |
| PreBackbone=Qwen2.5-7B Instruct2025.12 | 38.7 | |
| EtConBackbone=Qwen2.5-7B Instruct2025.12 | 38.4 | |
| FT-M+Con.Backbone=Qwen2.5-7B Instruct, Consolidation=true2025.12 | 37.8 | |
| AceMADBase Model=GPT-4o-mini, Iteration Rounds (T)=3, Number of Agents (N)=52026.03 | 37.74 | |
| FT-MBackbone=Qwen2.5-7B Instruct, Consolidation=false2025.12 | 37 | |
| ConSA (head-wise, single-layer)Target Sparsity (ρ)=0.50, Granularity=head-wise, Constraint Scope=single-layer, Model=1.7B2026.06 | 36.92 | |
| Qwen2.5-7BModel Type=Dense, Active Parameters=7B, Total Parameters=7B2026.02 | 36.4 | |
| ConSA (head-wise, all-layers)Target Sparsity (ρ)=0.50, Granularity=head-wise, Constraint Scope=all-layers, Model=1.7B2026.06 | 35.08 | |
| Rule (head-wise)Target Sparsity (ρ)=0.50, Granularity=head-wise, Model=1.7B2026.06 | 34.62 | |
| Dense FATarget Sparsity (ρ)=0, Model=1.7B2026.06 | 34.31 | |
| Qwen2.5-3BModel Type=Dense, Active Parameters=3B, Total Parameters=3B2026.02 | 33.5 | |
| ConSA (layer-wise)Target Sparsity (ρ)=0.50, Granularity=layer-wise, Model=1.7B2026.06 | 32.92 | |
| Qwen2.5-1.5BModel Type=Dense, Active Parameters=1.5B, Total Parameters=1.5B2026.02 | 31.9 | |
| Rule (layer-wise)Target Sparsity (ρ)=0.50, Granularity=layer-wise, Model=1.7B2026.06 | 31.71 | |
| Qwen2.5-1.5B#Params=1.5B/1.5B, #Tokens=18T2026.02 | 31.5 | |
| Qwen3-1.7B#Params=1.7B/1.7B, #Tokens=36T2026.02 | 31.5 | |
| LLaMA-MoE-v2-3.5BModel Type=MoE, Active Parameters=3.5B2026.02 | 30.7 | |
| LLaMA-MoE-3.0B#Params=3.0B/7B, #Tokens=2.2T2026.02 | 30.6 | |
| Llama-3.2-3BModel Type=Dense, Active Parameters=3B, Total Parameters=3B2026.02 | 30.6 | |
| SPES-9B#Params=3.1B/9B, #Tokens=400B2026.02 | 30.4 | |
| SmolLM2-1.7BModel Type=Dense, Active Parameters=1.7B, Total Parameters=1.7B2026.02 | 30.1 | |
| SmolLM2-1.7B#Params=1.7B/1.7B, #Tokens=11T2026.02 | 29.8 | |
| LLaMA-MoE-v1-3.5BModel Type=MoE, Active Parameters=3.5B2026.02 | 29.7 | |
| Qwen2.5-0.5B#Params=0.5B/0.5B, #Tokens=18T2026.02 | 29.5 | |
| OLMoE-1B-7BModel Type=MoE, Active Parameters=1B, Total Parameters=7B2026.02 | 29.3 | |
| Qwen3-0.6B#Params=0.6B/0.6B, #Tokens=36T2026.02 | 29 | |
| ExpertWeaver-E64-A14-S2(Qwen2.5-7B)Model Type=MoE, Active Parameters=3.5B, Total Parameters=7B, Training Budget=200B tokens2026.02 | 29 | |
| OLMoE-1B-7B*Model Type=MoE, Active Parameters=1B, Total Parameters=7B, Training Budget=500B tokens2026.02 | 28.7 | |
| ExpertWeaver-E64-A14-S2(OLMo-7B)Model Type=MoE, Active Parameters=1B, Total Parameters=7B, Training Budget=200B tokens2026.02 | 28.5 | |
| AceMADBase Model=Qwen3-235B-A22B-Instruct, Iteration Rounds (T)=5, Number of Agents (N)=52026.03 | 28.44 | |
| OLMoE-1B-7B#Params=1.3B/7B, #Tokens=5T2026.02 | 28.4 | |
| MoE++ 7B#Params=1.2B/7B, #Tokens=1T2026.02 | 28.4 | |
| Sheared-LLaMA-2.7BModel Type=Dense, Active Parameters=2.7B, Total Parameters=2.7B2026.02 | 28.3 | |
| Pythia-2.8BModel Type=Dense, Active Parameters=2.8B, Total Parameters=2.8B2026.02 | 28.1 | |
| Open-LLaMA-3B-v2Model Type=Dense, Active Parameters=3B, Total Parameters=3B2026.02 | 28.1 | |
| Pythia-2.8B#Params=2.8B/2.8B, #Tokens=300B2026.02 | 28 | |
| SPES-7B#Params=1.6B/7B, #Tokens=500B2026.02 | 27.5 | |
| OLMo-7BModel Type=Dense, Active Parameters=7B, Total Parameters=7B2026.02 | 27.5 | |
| INCITE-Base-3BModel Type=Dense, Active Parameters=3B, Total Parameters=3B2026.02 | 27.5 | |
| Pythia-1.4B#Params=1.4B/1.4B, #Tokens=300B2026.02 | 27.3 | |
| SPES-2B#Params=0.8B/2.1B, #Tokens=500B2026.02 | 27.2 | |
| OPT-1.3B#Params=1.3B/1.3B, #Tokens=180B2026.02 | 26.9 | |
| TinyLlama-1.1B#Params=1.1B/1.1B, #Tokens=3T2026.02 | 26.3 | |
| OPT-2.7B#Params=2.7B/2.7B, #Tokens=180B2026.02 | 26 | |
| OPT-2.7BModel Type=Dense, Active Parameters=2.7B, Total Parameters=2.7B2026.02 | 25.8 | |
| SOLARTraining Stage=SFT-1B2025.12 | 23.81 | |
| OpT-DeUSTraining Stage=SFT-1B2025.12 | 23.81 | |
| Avg-DeUSTraining Stage=SFT-1B2025.12 | 23.66 | |
| OpT-DeUSTraining Stage=CPT-1B2025.12 | 23.35 | |
| LESATraining Stage=SFT-1B2025.12 | 23.35 | |
| MIDUS-HMLTraining Stage=CPT-1B2025.12 | 23.2 | |
| Llama ProTraining Stage=SFT-1B2025.12 | 22.89 | |
| Avg-DeUSTraining Stage=CPT-1B2025.12 | 22.73 | |
| Gemma-2-2bModel Type=Dense, Active Parameters=2B, Total Parameters=2B2026.02 | 22.7 | |
| SOLARTraining Stage=CPT-1B2025.12 | 22.58 | |
| Llama ProTraining Stage=CPT-1B2025.12 | 22.43 | |
| BaseTraining Stage=CPT-1B2025.12 | 22.27 | |
| Decentralized MADBase Model=GPT-4o-mini, Number of Agents (N)=52026.03 | 22.19 | |
| Centralized MADBase Model=GPT-4o-mini, Number of Agents (N)=52026.03 | 22.19 | |
| Sparse MADBase Model=GPT-4o-mini, Number of Agents (N)=52026.03 | 22.19 | |
| LESATraining Stage=CPT-1B2025.12 | 21.97 | |
| Majority VotingBase Model=GPT-4o-mini, Number of Agents (N)=52026.03 | 21.88 | |
| ALPHABackbone=Qwen2.5-7B Instruct, Consolidation=false2025.12 | 21.8 |