General Knowledge Reasoning on MMLU Pro (Accuracy)
82.93AccuracyQwen3-Next-80B-A3B-Instruct
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3-Next-80B-A3B-Instruct2026.03 | 82.93 | |
| ParaManager-SFTOrchestration Strategy=SFT2026.04 | 81.43 | |
| ParaManagerOrchestration Strategy=Unified2026.04 | 81.43 | |
| PuppeteerOrchestration Strategy=Serial Orchestration2026.04 | 80.54 | |
| Qwen3-Coder 30B-A3B-InstructOrchestration Strategy=Self Refine2026.04 | 80.36 | |
| GPT-OSS-20BOrchestration Strategy=Base+Tool2026.04 | 80 | |
| GPT-OSS-20BOrchestration Strategy=Majority Vote2026.04 | 80 | |
| GPT-OSS-20BOrchestration Strategy=Self Refine2026.04 | 80 | |
| Qwen3-Coder 30B-A3B-InstructOrchestration Strategy=Majority Vote2026.04 | 80 | |
| RouterOrchestration Strategy=Static Workflow2026.04 | 80 | |
| ParaManager-SerialOrchestration Strategy=Serial2026.04 | 80 | |
| Qwen3-Omni-A3B-Instruct2026.03 | 79.89 | |
| EvoFlowOrchestration Strategy=Static Workflow2026.04 | 79.82 | |
| GPT-OSS-20BOrchestration Strategy=Base2026.04 | 79.64 | |
| ParaManager-MonoOrchestration Strategy=Mono2026.04 | 79.64 | |
| Meta Agent SearchOrchestration Strategy=Static Workflow2026.04 | 79.46 | |
| ToolOrchestraOrchestration Strategy=Serial Orchestration2026.04 | 79.46 | |
| Qwen3-30B A3B-Instruct-2507Orchestration Strategy=Majority Vote2026.04 | 78.57 | |
| Nemotron 3 Nano 30B-A3BOpen-Source=✓, Size=30B-A3B2026.04 | 78.3 | |
| Qwen3-30B A3B-Instruct-2507Orchestration Strategy=Base+Tool2026.04 | 77.86 | |
| Nemotron 3 Nano OmniOpen-Source=✓, Size=30B-A3B2026.04 | 77.3 | |
| Qwen3-Coder 30B-A3B-InstructOrchestration Strategy=Base2026.04 | 77.14 | |
| LongCat-Next2026.03 | 77.02 | |
| Qwen3-30B A3B-Instruct-2507Orchestration Strategy=Base2026.04 | 76.96 | |
| Qwen3-30B A3B-Instruct-2507Orchestration Strategy=Self Refine2026.04 | 76.79 | |
| Qwen3-4B Instruct-2507Orchestration Strategy=Self Refine2026.04 | 75.71 | |
| Qwen3-Coder 30B-A3B-InstructOrchestration Strategy=Base+Tool2026.04 | 75.71 | |
| Qwen3-4B Instruct-2507Orchestration Strategy=Base2026.04 | 74.82 | |
| Qwen3-4B Instruct-2507Orchestration Strategy=Base+Tool2026.04 | 74.11 | |
| Qwen3-4B Instruct-2507Orchestration Strategy=Majority Vote2026.04 | 72.86 | |
| Kimi-Linear-48B-A3B2026.03 | 67.22 | |
| HeRLModel=Qwen3-4B-Instruct-2507, Training Method=HeRL2026.03 | 62.1 | |
| Qwen3-OmniOpen-Source=✓, Size=30B-A3B2026.04 | 61.6 | |
| Qwen3-4B-Instruct-2507Model=Qwen3-4B-Instruct-2507, Training Method=Instruct2026.03 | 61.5 | |
| INTUITORBackbone=Qwen3-14B, Evaluation Protocol=standard evaluation protocol2025.05 | 61.3 | |
| Qwen3-14BBackbone=Qwen3-14B, Evaluation Protocol=standard evaluation protocol2025.05 | 59.7 | |
| GRPOBackbone=Qwen3-8B-Base2026.05 | 59.05 | |
| INTUITORBackbone=Qwen2.5-14B, Evaluation Protocol=standard evaluation protocol2025.05 | 58.3 | |
| SPIRALBackbone=Qwen3-8B-Base2026.05 | 58.28 | |
| STRATAGEMModel=STRATAGEM (Ours)2026.04 | 57.83 | |
| GRPOBackbone=Qwen2.5-14B, Evaluation Protocol=standard evaluation protocol2025.05 | 57.8 | |
| MARSBackbone=Qwen3-8B-Base2026.05 | 57.8 | |
| DEPTBackbone=Qwen3-8B-Base2026.05 | 57.6 | |
| Qwen2.5-7B-InstructModel=Qwen2.5-7B-Instruct, Training Method=Instruct2026.03 | 57.3 | |
| Qwen2.5-14BBackbone=Qwen2.5-14B, Evaluation Protocol=standard evaluation protocol2025.05 | 56.5 | |
| DEPTBackbone=Qwen3-4B-Base2026.05 | 56.45 | |
| HeRLModel=Qwen2.5-7B-Instruct, Training Method=HeRL2026.03 | 55.5 | |
| SPIRALModel=SPIRAL2026.04 | 53.93 | |
| MARSBackbone=Qwen3-4B-Base2026.05 | 53.34 | |
| SPAGBackbone=Qwen3-4B-Base2026.05 | 52.71 | |
| SPAGBackbone=Qwen3-8B-Base2026.05 | 51.46 | |
| INTUITORBackbone=Qwen2.5-7B, Evaluation Protocol=standard evaluation protocol2025.05 | 51.4 | |
| SPIRALBackbone=Qwen3-4B-Base2026.05 | 51.32 | |
| GRPOBackbone=Qwen2.5-7B, Evaluation Protocol=standard evaluation protocol2025.05 | 51.1 | |
| PSFTTraining Protocol=PSFT2025.08 | 50.55 | |
| Qwen2.5-7BBackbone=Qwen2.5-7B, Evaluation Protocol=standard evaluation protocol2025.05 | 49.7 | |
| GRPOBackbone=Qwen3-4B-Base2026.05 | 49.14 | |
| BaseTraining Protocol=Base2025.08 | 47.9 | |
| Qwen3-4B-BaseModel=Qwen3-4B-Base2026.04 | 47.2 | |
| VANILLABackbone=Qwen3-8B-Base2026.05 | 46.97 | |
| SFTTraining Protocol=SFT2025.08 | 40.67 | |
| VANILLABackbone=Qwen3-4B-Base2026.05 | 39.36 | |
| HeRLModel=Llama-3.2-3B-Instruct, Training Method=HeRL2026.03 | 34.8 | |
| Llama-3.2-3B-InstructModel=Llama-3.2-3B-Instruct, Training Method=Instruct2026.03 | 33.9 |