General Knowledge on MMLU (Accuracy)
91.2MMLU General Knowledge AccuracyQwen3.5-35B-A3B
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3.5-35B-A3B#Token=9002026.04 | 91.2 | |
| DeepSeek V3.2Evaluation Mode=Chat2025.12 | 91.1 | |
| GLM 4.6Evaluation Mode=Chat2025.12 | 90.7 | |
| LongCat-Flash ChatEvaluation Mode=Chat2025.12 | 89.7 | |
| LongCat-Flash Exp-ChatEvaluation Mode=Chat2025.12 | 89.6 | |
| JoyAI-LLM Flash#Token=2002026.04 | 89.1 | |
| Qwen3.5-35B-A3B-BaseEvaluation Framework=OpenCompass, Decoding=Greedy2026.04 | 88.4 | |
| GLM-4.7-Flash-T#Token=20002026.04 | 88.2 | |
| Qwen3-30B-A3B#Token=2002026.04 | 88 | |
| Kimi-K2 Base# Shots=5-shot, # Activated Params=32B, # Total Params=1043B2026.01 | 87.8 | |
| DeepSeek-V3.2 Exp Base# Shots=5-shot, # Activated Params=37B, # Total Params=671B2026.01 | 87.8 | |
| DenserBackbone=Qwen3-32B-think2025.12 | 87.8 | |
| Process SupervisionBackbone=Qwen3-32B-think2025.12 | 87.7 | |
| DeepSeek-V3.1 Base# Shots=5-shot, # Activated Params=37B, # Total Params=671B2026.01 | 87.4 | |
| Reflection-CoTBackbone=Qwen3-32B-think2025.12 | 87.3 | |
| MiMo-V2-Flash Base# Shots=5-shot, # Activated Params=15B, # Total Params=309B2026.01 | 86.7 | |
| Self-VerificationBackbone=Qwen3-32B-think2025.12 | 86.4 | |
| Qwen3-Next-80B-A3B#Token=4002026.04 | 86.4 | |
| DenserBackbone=Qwen3-32B2025.12 | 86.2 | |
| Tree-of-ThoughtBackbone=Qwen3-32B-think2025.12 | 86.1 | |
| N-3-Super 120B-A12B-BaseShots=52026.04 | 86.01 | |
| Process SupervisionBackbone=Qwen3-32B2025.12 | 85.9 | |
| Self-ConsistencyBackbone=Qwen3-32B-think2025.12 | 85.6 | |
| Reflection-CoTBackbone=Qwen3-32B2025.12 | 85.5 | |
| Qwen-3-30B-A3BOpenness=Open-weights, Regional Origin=Non-European, Post-training=Instruction-tuned2026.02 | 85 | |
| Think-to-ThinkBackbone=Qwen3-32B-think2025.12 | 84.9 | |
| Self-VerificationBackbone=Qwen3-32B2025.12 | 84.7 | |
| JoyAI-LLM Flash-BaseEvaluation Framework=OpenCompass, Decoding=Greedy2026.04 | 84.7 | |
| Llama-3.3-70BOpenness=Open-weights, Regional Origin=Non-European, Post-training=Instruction-tuned2026.02 | 84.6 | |
| Tree-of-ThoughtBackbone=Qwen3-32B2025.12 | 84.3 | |
| Qwen-3-32BOpenness=Open-weights, Regional Origin=Non-European, Post-training=Instruction-tuned2026.02 | 84 | |
| Self-ConsistencyBackbone=Qwen3-32B2025.12 | 83.8 | |
| Qwen3Parameters=8B, Architecture=AR2026.04 | 83.5 | |
| Chain-of-ThoughtBackbone=Qwen3-32B-think2025.12 | 83.2 | |
| Think-to-ThinkBackbone=Qwen3-32B2025.12 | 83.1 | |
| NBDiffParameters=7B2026.04 | 82.9 | |
| I-DLMParameters=8B2026.04 | 82.4 | |
| Qwen3-30B-A3B-BaseEvaluation Framework=OpenCompass, Decoding=Greedy2026.04 | 82.1 | |
| DenserBackbone=Qwen3-8B-think2025.12 | 81.8 | |
| Qwen3-30B-A3BSparsity Level=Base, K=82026.05 | 81.8 | |
| Process SupervisionBackbone=Qwen3-8B-think2025.12 | 81.7 | |
| Chain-of-ThoughtBackbone=Qwen3-32B2025.12 | 81.5 | |
| Reflection-CoTBackbone=Qwen3-8B-think2025.12 | 81.3 | |
| Qwen-3-14BOpenness=Open-weights, Regional Origin=Non-European, Post-training=Instruction-tuned2026.02 | 81.2 | |
| Ling-flash base-2.0Shots=52026.04 | 81 | |
| GLM-4.5 Air-BaseShots=52026.04 | 81 | |
| Gemma-3-27BOpenness=Open-weights, Regional Origin=Non-European, Post-training=Instruction-tuned2026.02 | 80.4 | |
| DenserBackbone=Qwen3-14B-no-think2025.12 | 80.3 | |
| DenserBackbone=Qwen3-8B2025.12 | 80.2 | |
| Process SupervisionBackbone=Qwen3-14B-no-think2025.12 | 80.2 | |
| OLMo-3.1-32BOpenness=Fully-open, Regional Origin=Non-European, Post-training=Instruction-tuned2026.02 | 80.1 | |
| Tree-of-ThoughtBackbone=Qwen3-8B-think2025.12 | 80.1 | |
| Self-VerificationBackbone=Qwen3-8B-think2025.12 | 80.1 | |
| BEAMSparsity Level=Mid Sparsity, beta=0.012026.05 | 80.09 | |
| Process SupervisionBackbone=Qwen3-8B2025.12 | 80 | |
| Reflection-CoTBackbone=Qwen3-14B-no-think2025.12 | 79.8 | |
| Reflection-CoTBackbone=Qwen3-8B2025.12 | 79.6 | |
| Self-ConsistencyBackbone=Qwen3-8B-think2025.12 | 79.6 | |
| DenserBackbone=Qwen3-4B-think2025.12 | 79.1 | |
| Think-to-ThinkBackbone=Qwen3-8B-think2025.12 | 78.9 | |
| Self-VerificationBackbone=Qwen3-14B-no-think2025.12 | 78.7 | |
| Process SupervisionBackbone=Qwen3-4B-think2025.12 | 78.6 | |
| Self-VerificationBackbone=Qwen3-8B2025.12 | 78.6 | |
| Tree-of-ThoughtBackbone=Qwen3-14B-no-think2025.12 | 78.6 | |
| SDAR-8BScale=8B, Tokens=55B, FLOPs=2640, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 78.6 | |
| Tree-of-ThoughtBackbone=Qwen3-8B2025.12 | 78.4 | |
| Top-K ReducedSparsity Level=Mid Sparsity, K=42026.05 | 78.27 | |
| Reflection-CoTBackbone=Qwen3-4B-think2025.12 | 78.2 | |
| Self-ConsistencyBackbone=Qwen3-14B-no-think2025.12 | 78.1 | |
| Self-ConsistencyBackbone=Qwen3-8B2025.12 | 77.9 | |
| MoE-DynamicSparsity Level=Mid Sparsity, phi=0.32026.05 | 77.86 | |
| Self-VerificationBackbone=Qwen3-4B-think2025.12 | 77.5 | |
| Chain-of-ThoughtBackbone=Qwen3-8B-think2025.12 | 77.4 | |
| Think-to-ThinkBackbone=Qwen3-14B-no-think2025.12 | 77.4 | |
| Mistral-3.2-24BOpenness=Open-weights, Regional Origin=European, Post-training=Instruction-tuned2026.02 | 77.3 | |
| Think-to-ThinkBackbone=Qwen3-8B2025.12 | 77.2 | |
| Tree-of-ThoughtBackbone=Qwen3-4B-think2025.12 | 77 | |
| Qwen3-8BGen. Mode=AR, Avg TPF=1.002026.07 | 76.93 | |
| DenserBackbone=Qwen3-4B2025.12 | 76.9 | |
| Qwen3-8BParameters=8B2026.01 | 76.8 | |
| Molmo2-8BParameters=8B2026.01 | 76.6 | |
| TiDARParameters=8B2026.04 | 76.6 | |
| Self-ConsistencyBackbone=Qwen3-4B-think2025.12 | 76.5 | |
| Ministral3-8BGen. Mode=AR, Avg TPF=1.002026.07 | 76.39 | |
| Process SupervisionBackbone=Qwen3-4B2025.12 | 76.2 | |
| Gemma-3-12BOpenness=Open-weights, Regional Origin=Non-European, Post-training=Instruction-tuned2026.02 | 76.1 | |
| BEAMSparsity Level=High Sparsity, beta=0.12026.05 | 76.06 | |
| Chain-of-ThoughtBackbone=Qwen3-14B-no-think2025.12 | 75.9 | |
| Reflection-CoTBackbone=Qwen3-4B2025.12 | 75.8 | |
| Think-to-ThinkBackbone=Qwen3-4B-think2025.12 | 75.8 | |
| Chain-of-ThoughtBackbone=Qwen3-8B2025.12 | 75.8 | |
| Qwen3-4BParams=4B2025.12 | 75.78 | |
| WeDLMParameters=8B2026.04 | 75.5 | |
| BaseAllowed Domains=—2026.05 | 75.4 | |
| Self-VerificationBackbone=Qwen3-4B2025.12 | 75.2 | |
| SDAR-4BScale=4B, Tokens=55B, FLOPs=1320, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 74.9 | |
| PALETTEAllowed Domains=Illegal | Violence2026.05 | 74.8 | |
| Tree-of-ThoughtBackbone=Qwen3-4B2025.12 | 74.7 | |
| Nemotron-Labs-Diffusion-8BGen. Mode=AR, Avg TPF=1.002026.07 | 74.68 | |
| Nemotron-Labs-Diffusion-8BGen. Mode=Diff., Avg TPF=2.062026.07 | 74.68 |