Mathematical Reasoning on MATH 500 (Accuracy, Cost, Latency)
98AccuracyQwen3-Next-80B-A3B-Instruct
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Qwen3-Next-80B-A3B-InstructArchitecture=MoE, # Total Params=80B, # Activated Params=3B2026.01 | 98 | — | — | |
| BF16Model=Nemotron Nano V22026.01 | 97.8 | — | — | |
| NVFP4 PTQModel=Nemotron Nano V22026.01 | 97.2 | — | — | |
| NVFP4 QATModel=Nemotron Nano V22026.01 | 97.2 | — | — | |
| NVFP4 QADModel=Nemotron Nano V22026.01 | 97.2 | — | — | |
| LongCat-Flash-LiteArchitecture=MoE + NE, # Total Params=68.5B, # Activated Params=2.9B~4.5B2026.01 | 96.8 | — | — | |
| Critique-GRPOCritique Model=GPT-4o, Decoding budget=16,384 tokens2025.06 | 96.4 | — | — | |
| BF16Model=Llama Nemotron Super V12026.01 | 95.8 | — | — | |
| Gemini 2.5 Flash-Lite2026.01 | 95.2 | — | — | |
| NVFP4 QADModel=Llama Nemotron Super V12026.01 | 94.6 | — | — | |
| NVFP4 QATModel=Llama Nemotron Super V12026.01 | 94.3 | — | — | |
| Kimi-Linear-48B-A3BArchitecture=MoE, # Total Params=48B, # Activated Params=3B2026.01 | 94.2 | — | — | |
| R1-GRPOCritique Model=None, Decoding budget=16,384 tokens2025.06 | 93 | — | — | |
| Qwen3-32BCritique Model=None, Decoding budget=16,384 tokens2025.06 | 91.8 | — | — | |
| NVFP4 PTQModel=Llama Nemotron Super V12026.01 | 91.4 | — | — | |
| OracleType=Single LLM with Routing2026.01 | 83.8 | — | — | |
| Qwen2.5-Math-7B-InstructType=Single LLM2026.01 | 80.7 | — | — | |
| RouteMoAType=Ours2026.01 | 76 | 4.03 | 19.05 | |
| MoAType=Multi-LLMs2026.01 | 73.6 | 19.68 | 26.62 | |
| SMoAType=Multi-LLMs2026.01 | 73.5 | 4.4 | 23.16 | |
| RouterDCType=Single LLM with Routing2026.01 | 72.8 | — | — | |
| Qwen2.5-Coder-7B-InstructType=Single LLM2026.01 | 65.3 | — | — | |
| RouteLLMType=Single LLM with Routing2026.01 | 64.3 | — | — | |
| Ministral-8B-Instruct-2410Type=Single LLM2026.01 | 51 | — | — | |
| Gemma-2-9B-itType=Single LLM2026.01 | 46.5 | — | — | |
| Bio-Medical-Llama-3-8BType=Single LLM2026.01 | 11.7 | — | — |