General Reasoning on MMLU
95.1MMLU AccuracyM2CL
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| M2CLBackbone Model=Qwen-72B, Number of LLMs=42026.02 | 95.1 | — | |
| M2CLBackbone Model=Qwen-14B, Number of LLMs=42026.02 | 93.7 | — | |
| M2CLBackbone Model=Qwen-7B, Number of LLMs=42026.02 | 92.5 | — | |
| DyLANBackbone Model=Qwen-72B, Number of LLMs=42026.02 | 91.5 | — | |
| GPTSwarmBackbone Model=Qwen-72B, Number of LLMs=42026.02 | 91.5 | — | |
| Process SupervisionModel=GPT-4o2025.12 | 89 | — | |
| DenserModel=GPT-4o2025.12 | 88.7 | — | |
| Reflection-CoTModel=GPT-4o2025.12 | 88.6 | — | |
| Latent-GRPOModel Scale=Qwen3-4B2026.01 | 88.5 | 3,108.44 | |
| RAPSNumber of agents=52026.02 | 88.2 | — | |
| Tree-of-ThoughtModel=GPT-4o2025.12 | 87.9 | — | |
| Self-VerificationModel=GPT-4o2025.12 | 87.9 | — | |
| Process SupervisionModel=DeepSeek-V32025.12 | 87.8 | — | |
| Self-ConsistencyModel=GPT-4o2025.12 | 87.4 | — | |
| Reflection-CoTModel=DeepSeek-V32025.12 | 87.4 | — | |
| DenserModel=DeepSeek-V32025.12 | 87.1 | — | |
| GPTSwarmBackbone Model=Qwen-14B, Number of LLMs=42026.02 | 87 | — | |
| DyLANBackbone Model=Qwen-14B, Number of LLMs=42026.02 | 86.8 | — | |
| Think-to-ThinkModel=GPT-4o2025.12 | 86.7 | — | |
| G-DesignerNumber of agents=52026.02 | 86.3 | — | |
| Tree-of-ThoughtModel=DeepSeek-V32025.12 | 86.2 | — | |
| Self-VerificationModel=DeepSeek-V32025.12 | 86.2 | — | |
| InfiFusion*Model Size=14B, GPU Hours=160, Subset Source Models=true2025.05 | 85.81 | — | |
| Chain-of-ThoughtModel=GPT-4o2025.12 | 85.7 | — | |
| Self-ConsistencyModel=DeepSeek-V32025.12 | 85.7 | — | |
| Phi-4Model Size=14B, GPU Hours=∼1.0M2025.05 | 85.62 | — | |
| RandomNumber of agents=52026.02 | 85.6 | — | |
| AFlowNumber of agents=52026.02 | 85.6 | — | |
| GRPO (LLM-Judge)Model Scale=Qwen3-4B2026.01 | 85.5 | 6,753.47 | |
| LLM-DebateNumber of agents=52026.02 | 85 | — | |
| MaASNumber of agents=52026.02 | 85 | — | |
| Think-to-ThinkModel=DeepSeek-V32025.12 | 85 | — | |
| SFTModel Size=14B, GPU Hours=152025.05 | 84.36 | — | |
| ChainNumber of agents=52026.02 | 84.3 | — | |
| AgentPruneNumber of agents=52026.02 | 84.3 | — | |
| PuppeteerNumber of agents=52026.02 | 84.3 | — | |
| Process SupervisionModel=o1-mini2025.12 | 84.3 | — | |
| SFT-DPOModel Size=14B, GPU Hours=502025.05 | 84.27 | — | |
| InfiFPOModel Size=14B, GPU Hours=582025.05 | 84.27 | — | |
| FuseChat*Model Size=14B, GPU Hours=650, Subset Source Models=true2025.05 | 84.23 | — | |
| Best-of-N (BoN)Backbone Model=Qwen-72B, Number of LLMs=42026.02 | 84.2 | — | |
| SFT-IPOModel Size=14B, GPU Hours=502025.05 | 84.08 | — | |
| DenserModel=o1-mini2025.12 | 84 | — | |
| SFT-WRPOModel Size=14B, GPU Hours=572025.05 | 83.98 | — | |
| FuseLLM*Model Size=14B, GPU Hours=225, Subset Source Models=true2025.05 | 83.92 | — | |
| Reflection-CoTModel=o1-mini2025.12 | 83.9 | — | |
| MacNetBackbone Model=Qwen-72B, Number of LLMs=42026.02 | 83.8 | — | |
| Chain-of-ThoughtModel=DeepSeek-V32025.12 | 83.8 | — | |
| ComplexCoTNumber of agents=12026.02 | 83.7 | — | |
| GPTSwarmNumber of agents=52026.02 | 83.7 | — | |
| BaseModel Scale=Qwen3-4B2026.01 | 83.7 | — | |
| InfiFPO*Model Size=14B, GPU Hours=55, Subset Source Models=true2025.05 | 83.33 | — | |
| DiSCTTModel=Qwen-2.5-7B-Instruct2026.03 | 83.3 | — | |
| CoTNumber of agents=12026.02 | 83 | — | |
| MAS-ZeroNumber of agents=52026.02 | 83 | — | |
| Self-VerificationModel=o1-mini2025.12 | 82.8 | — | |
| DebateBackbone Model=Qwen-72B, Number of LLMs=42026.02 | 82.7 | — | |
| Tree-of-ThoughtModel=o1-mini2025.12 | 82.6 | — | |
| SCNumber of agents=12026.02 | 82.4 | — | |
| TreeNumber of agents=52026.02 | 82.4 | — | |
| AutoAgentsNumber of agents=52026.02 | 82.4 | — | |
| Self-ConsistencyModel=o1-mini2025.12 | 82.1 | — | |
| Vanilla IONumber of agents=1, Backbone=GPT-4o-mini2026.02 | 81.7 | — | |
| Mistral-SmallModel Size=24B, GPU Hours=∼1.6M2025.05 | 81.69 | — | |
| Think-to-ThinkModel=o1-mini2025.12 | 81.4 | — | |
| DiSCTTModel=Qwen-3-4B-Base2026.03 | 81.3 | — | |
| LLM-BlenderNumber of agents=52026.02 | 81 | — | |
| StarNumber of agents=52026.02 | 80.4 | — | |
| Qwen2.5-InstructModel Size=14B, GPU Hours=∼1.8M2025.05 | 80.22 | — | |
| Chain-of-ThoughtModel=o1-mini2025.12 | 80.2 | — | |
| Best-of-N (BoN)Backbone Model=Qwen-14B, Number of LLMs=42026.02 | 79.7 | — | |
| SDAR-30B-A3Bgeneration length=optimal, block length=optimal, denoising steps=optimal2026.04 | 79.16 | — | |
| MacNetBackbone Model=Qwen-14B, Number of LLMs=42026.02 | 78.9 | — | |
| Qwen2.5-7B-InstructSetting=Teacher2026.05 | 78.22 | — | |
| EVOL-RLModel=Qwen-2.5-7B-Instruct2026.03 | 77.9 | — | |
| Gemma-3-InstructModel Size=12B2025.05 | 77.61 | — | |
| SDAR-8B-Chatgeneration length=optimal, block length=optimal, denoising steps=optimal2026.04 | 77.23 | — | |
| DebateBackbone Model=Qwen-14B, Number of LLMs=42026.02 | 77.2 | — | |
| Baseline (full model)Sparsity=0%, Backbone=Phi-4-14B2025.05 | 77.06 | — | |
| TEALSparsity=25%, Backbone=Phi-4-14B2025.05 | 76.63 | — | |
| WINASparsity=25%, Backbone=Phi-4-14B2025.05 | 76.6 | — | |
| WINASparsity=40%, Backbone=Phi-4-14B2025.05 | 76.44 | — | |
| TTRLModel=Qwen-2.5-7B-Instruct2026.03 | 76.4 | — | |
| GPTSwarmBackbone Model=Qwen-7B, Number of LLMs=42026.02 | 76.3 | — | |
| BaseModel=Qwen-2.5-7B-Instruct2026.03 | 76.2 | — | |
| WINASparsity=50%, Backbone=Phi-4-14B2025.05 | 75.83 | — | |
| TEALSparsity=40%, Backbone=Phi-4-14B2025.05 | 75.1 | — | |
| Qwen2.5-CoderModel Size=14B, GPU Hours=∼1.8M2025.05 | 75.08 | — | |
| EVOL-RLModel=Qwen-3-4B-Base2026.03 | 74.5 | — | |
| DyLANBackbone Model=Qwen-7B, Number of LLMs=42026.02 | 74.3 | — | |
| Best-of-N (BoN)Backbone Model=Qwen-7B, Number of LLMs=42026.02 | 74.2 | — | |
| TEALSparsity=50%, Backbone=Phi-4-14B2025.05 | 73.52 | — | |
| LLaDA2.0-minigeneration length=optimal, block length=optimal, denoising steps=optimal2026.04 | 72.54 | — | |
| Single executionBackbone Model=Qwen-72B, Number of LLMs=42026.02 | 72.5 | — | |
| TTRLModel=Qwen-3-4B-Base2026.03 | 72.4 | — | |
| MacNetBackbone Model=Qwen-7B, Number of LLMs=42026.02 | 71.5 | — | |
| DebateBackbone Model=Qwen-7B, Number of LLMs=42026.02 | 71.1 | — | |
| GPTSwarmAggregation=AgentAuditor2026.02 | 70.75 | — | |
| Group-DebateAggregation=AgentAuditor2026.02 | 70.58 | — | |
| AgentPruneAggregation=AgentAuditor2026.02 | 70.12 | — |