Language Understanding on MMLU (Accuracy)
96.6AccuracyM2CL
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| M2CLModel=Qwen-72B, Number of LLMs=162026.02 | 96.6 | — | |
| M2CLModel=Qwen-14B, Number of LLMs=162026.02 | 95.9 | — | |
| M2CLBase Model=Llama-70B, Number of LLMs=42026.02 | 95.6 | — | |
| M2CLModel=Qwen-7B, Number of LLMs=162026.02 | 94.6 | — | |
| M2CLModel=Llama-14B, Number of LLMs=162026.02 | 94.5 | — | |
| Conductorparameters=7B2025.12 | 94.1 | — | |
| M2CLModel=Llama-70B, Number of LLMs=162026.02 | 93.9 | — | |
| M2CLBase Model=Qwen-72B, Number of LLMs=42026.02 | 93.5 | — | |
| GPT 52025.12 | 93.5 | — | |
| DyLANModel=Llama-70B, Number of LLMs=162026.02 | 93.2 | — | |
| GPTSwarmModel=Llama-70B, Number of LLMs=162026.02 | 93 | — | |
| DyLANModel=Qwen-72B, Number of LLMs=162026.02 | 92.7 | — | |
| GPTSwarmModel=Qwen-72B, Number of LLMs=162026.02 | 92.4 | — | |
| Gemini 2.5 Pro2025.12 | 92.4 | — | |
| M2CLBase Model=Qwen-14B, Number of LLMs=42026.02 | 91.4 | — | |
| Claude Sonnet 42025.12 | 91.4 | — | |
| Graph-GRPO2026.03 | 90.12 | — | |
| Absolute SOTA2024.06 | 90 | — | |
| Claude 3.5 SonnetEvaluation Protocol=0-shot2026.03 | 89.9 | — | |
| OFA-MAS (fine-tuned)Training protocol=One-for-All Topology Design (Fine-tuned)2026.01 | 89.54 | — | |
| GPT-4oEvaluation Protocol=0-shot2026.03 | 89.1 | — | |
| EIB-LEARNERTraining protocol=Dataset-Specific Training2026.01 | 88.9 | — | |
| OFA-MAS (pre-trained)Training protocol=One-for-All Topology Design (Pre-trained)2026.01 | 88.9 | — | |
| EIB-LEARNER2026.03 | 88.9 | — | |
| M2CLBase Model=Qwen-7B, Number of LLMs=42026.02 | 88.7 | — | |
| KALEBackbone=Qwen2.5 32B, Category=Augmented-based2026.01 | 88.59 | — | |
| MacNetModel=Qwen-72B, Number of LLMs=162026.02 | 88.1 | — | |
| DyLANModel=Qwen-14B, Number of LLMs=162026.02 | 88 | — | |
| GPTSwarmModel=Qwen-14B, Number of LLMs=162026.02 | 88 | — | |
| Llama 3 405BEvaluation Protocol=0-shot2026.03 | 87.3 | — | |
| DeepSeek-V3-Base#Shots=5-shot, Architecture=MoE, # activated params=37B, # total params=671B2026.01 | 87.1 | — | |
| G-designerTraining protocol=Dataset-Specific Training2026.01 | 86.92 | — | |
| G-designer2026.03 | 86.92 | — | |
| M2CLBase Model=Llama-14B, Number of LLMs=42026.02 | 86.9 | — | |
| MacNetModel=Llama-14B, Number of LLMs=162026.02 | 86.8 | — | |
| MacNetModel=Llama-70B, Number of LLMs=162026.02 | 86.7 | — | |
| GPT-4shot=5-shot2023.06 | 86.4 | — | |
| GPT-4Model Type=Pre-trained2024.07 | 86.4 | — | |
| Llama3.1-70B-InstructLength=128k2024.08 | 86 | — | |
| GPT3MixBackbone=Qwen2.5 32B, Category=Augmented-based2026.01 | 85.69 | — | |
| AgentDropoutTraining protocol=Dataset-Specific Training2026.01 | 85.62 | — | |
| AgentDropout2026.03 | 85.62 | — | |
| STaRBackbone=Qwen2.5 32B, Category=Augmented-based2026.01 | 85.24 | — | |
| Llama 3 405BModel Type=Pre-trained2024.07 | 85.2 | — | |
| DMTBackbone=Qwen2.5 32B, Category=SFT-based2026.01 | 85.17 | — | |
| AgentPruneTraining protocol=Dataset-Specific Training2026.01 | 85.07 | — | |
| AgentPrune2026.03 | 85.07 | — | |
| AugGPTBackbone=Qwen2.5 32B, Category=Augmented-based2026.01 | 85.04 | — | |
| LLM-DebateAgent topology=Fixed Multi-Agent (LLM-Debate)2026.01 | 84.96 | — | |
| LLM-Debate2026.03 | 84.96 | — | |
| GPTSwarmBase Model=Qwen-72B, Number of LLMs=42026.02 | 84.9 | — | |
| GraphRAGBackbone=Qwen2.5 32B, Category=Retrieval-based2026.01 | 84.85 | — | |
| DyLANBase Model=Qwen-72B, Number of LLMs=42026.02 | 84.6 | — | |
| LLaMA-3.1-405B Base#Shots=5-shot, Architecture=Dense, # activated params=405B, # total params=405B2026.01 | 84.4 | — | |
| R1-Distill-Qwen-32B2025.12 | 84.4 | — | |
| RandomAgent topology=Fixed Multi-Agent (Random)2026.01 | 84.31 | — | |
| Random2026.03 | 84.31 | — | |
| KG-SFTBackbone=Qwen2.5 32B, Category=SFT-based2026.01 | 84.26 | — | |
| Qwen2-72B2024.07 | 84.2 | — | |
| Qwen2-72B-InstructLength=128k2024.08 | 84.2 | — | |
| BoNModel=Qwen-72B, Number of LLMs=162026.02 | 84.2 | — | |
| BoNBase Model=Qwen-72B, Number of LLMs=42026.02 | 84.2 | — | |
| SDFTBackbone=Qwen2.5 32B, Category=SFT-based2026.01 | 84.13 | — | |
| Qwen3-32B (thinking)mode=thinking2025.12 | 84.1 | — | |
| DebateModel=Qwen-72B, Number of LLMs=162026.02 | 83.9 | — | |
| Gemini UltraModel Type=Pre-trained2024.07 | 83.7 | — | |
| SC (CoT)Prompting strategy=Single-Agent Prompting2026.01 | 83.66 | — | |
| SC (CoT)2026.03 | 83.66 | — | |
| MeanLearnBackbone=Qwen2.5 32B, Category=SFT-based2026.01 | 83.61 | — | |
| Llama 3 70BEvaluation Protocol=0-shot2026.03 | 83.6 | — | |
| Qwen3-32B2025.12 | 83.5 | — | |
| StructGPTBackbone=Qwen2.5 32B, Category=Retrieval-based2026.01 | 83.41 | — | |
| TOGBackbone=Qwen2.5 32B, Category=Retrieval-based2026.01 | 83.27 | — | |
| DyLANModel=Llama-14B, Number of LLMs=162026.02 | 83.2 | — | |
| MacNetModel=Qwen-14B, Number of LLMs=162026.02 | 83.1 | — | |
| ChainAgent topology=Fixed Multi-Agent (Chain)2026.01 | 83.01 | — | |
| Chain2026.03 | 83.01 | — | |
| BoNModel=Llama-70B, Number of LLMs=162026.02 | 83 | — | |
| BoNBase Model=Llama-70B, Number of LLMs=42026.02 | 83 | — | |
| DebateModel=Llama-70B, Number of LLMs=162026.02 | 82.9 | — | |
| NBDiff-7B-INSTRUCTParameters=7B, Training Protocol=Instruct, Sampling=Standard2025.12 | 82.9 | — | |
| SFTBackbone=Qwen2.5 32B, Category=SFT-based2026.01 | 82.82 | — | |
| Trinity Large Baseshot=5-shot2026.02 | 82.58 | — | |
| CompleteAgent topology=Fixed Multi-Agent (Complete)2026.01 | 82.35 | — | |
| Complete2026.03 | 82.35 | — | |
| DyLANBase Model=Llama-70B, Number of LLMs=42026.02 | 82.1 | — | |
| PlaintextBase Model=Qwen2.5-14B-Instruct2026.03 | 81.95 | — | |
| Gemini-1.5-ProLength=128k2024.08 | 81.9 | — | |
| CoTPrompting strategy=Single-Agent Prompting2026.01 | 81.69 | — | |
| CoT2026.03 | 81.69 | — | |
| CoTBackbone=Qwen2.5 32B, Category=Prompt-based2026.01 | 81.65 | — | |
| GPTSwarmBase Model=Llama-70B, Number of LLMs=42026.02 | 81.4 | — | |
| ORCHType=Multi-agent orchestration (3-agent + merge), Average latency (ms/question)=11,775, Description=Highest accuracy, highest cost and latency2026.02 | 81.3 | — | |
| gemma-3-27b-itinstruct=true2025.12 | 81.3 | — | |
| Nemotron 4 340BModel Type=Pre-trained2024.07 | 81.1 | — | |
| TreeAgent topology=Fixed Multi-Agent (Tree)2026.01 | 81.04 | — | |
| Tree2026.03 | 81.04 | — | |
| GPT-4Zero-shot=true2023.11 | 80.61 | — | |
| AloePriBase Model=Qwen2.5-14B-Instruct2026.03 | 80.61 | — | |
| GPT-4Length=128k2024.08 | 80.5 | — |