Language Understanding on CMMLU
90.1AccuracyQwen2-72B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen2-72B2024.07 | 90.1 | — | |
| Qwen2-57B-A14BArchitecture=MoE, # Act Params=14B, # Params=57B2024.07 | 88.5 | — | |
| Qwen1.5-110B2024.07 | 88.3 | — | |
| Yi-1.5-34BArchitecture=Dense, # Act Params=32B, # Params=32B2024.07 | 84.8 | — | |
| Qwen3-30B-A3BModel=Qwen3-30B-A3B, Target Top-K (K)=8, Avg. K=8.002026.05 | 83.69 | — | |
| Qwen1.5-72B2024.07 | 83.5 | — | |
| Qwen1.5-32BArchitecture=Dense, # Act Params=34B, # Params=34B2024.07 | 82.3 | — | |
| BEAMModel=Qwen3-30B-A3B, Beta (β)=0.01, Avg. K=4.232026.05 | 81.53 | — | |
| DyMoEModel=Qwen3-30B-A3B, High / Low Bits=4 / 0, Retention ratio (r)=1.02026.03 | 80.44 | — | |
| DyMoEModel=Qwen3-30B-A3B, High / Low Bits=4 / 2, Retention ratio (r)=1.02026.03 | 80.44 | — | |
| DyMoEModel=Qwen3-30B-A3B, High / Low Bits=4 / 0, Retention ratio (r)=0.92026.03 | 80.19 | — | |
| DyMoEModel=Qwen3-30B-A3B, High / Low Bits=4 / 2, Retention ratio (r)=0.92026.03 | 80.15 | — | |
| DyMoEModel=Qwen3-30B-A3B, High / Low Bits=4 / 0, Retention ratio (r)=0.752026.03 | 79.6 | — | |
| DyMoEModel=Qwen3-30B-A3B, High / Low Bits=4 / 2, Retention ratio (r)=0.752026.03 | 78.94 | — | |
| Qwen3-8BCategory=Naive LLM2026.01 | 76.8 | — | |
| Qwen1.5-MoEModel=Qwen1.5-MoE-A2.7B, Target Top-K (K)=4, Avg. K=4.002026.05 | 75.18 | — | |
| VIST2-8BCategory=Ours2026.01 | 75.12 | — | |
| Qwen3-VL-8BCategory=Visual-enhanced LLM2026.01 | 71.35 | — | |
| Qwen2-1.5B# Non-Emb Params=1.2B2024.07 | 70.3 | — | |
| Qwen3-4BCategory=Naive LLM2026.01 | 70.25 | — | |
| Qwen1.5-7BArchitecture=Dense, Decoding Framework=vLLM v0.8.4, Precision=bfloat16, max_model_len=4096, temperature=0.1, seed=12342026.06 | 69.6 | — | |
| VIST2-4BCategory=Ours2026.01 | 69.5 | — | |
| BEAMModel=Qwen1.5-MoE-A2.7B, Beta (β)=0.01, Avg. K=1.562026.05 | 69.18 | — | |
| Llama-3-70B2024.07 | 67.2 | — | |
| Qwen3-VL-4BCategory=Visual-enhanced LLM2026.01 | 65.42 | — | |
| MiniCPM4# Total Params=0.5B2026.05 | 65.22 | — | |
| Fisher-MoEArchitecture=Sparse MoE (compressed), Compression Ratio=50%, Calibration Method=128 GSM8K calibration samples, Decoding Framework=vLLM v0.8.4, Precision=bfloat16, max_model_len=4096, temperature=0.1, seed=12342026.06 | 61.6 | — | |
| Qwen3# Total Params=1.7B2026.05 | 61.22 | — | |
| DeepSeekV2-LiteModel=DeepSeekV2-Lite, Target Top-K (K)=6, Avg. K=6.002026.05 | 60.93 | — | |
| Qwen1.5-MoE-A2.7BArchitecture=Sparse MoE, Decoding Framework=vLLM v0.8.4, Precision=bfloat16, max_model_len=4096, temperature=0.1, seed=12342026.06 | 60.7 | — | |
| BEAMModel=DeepSeekV2-Lite, Beta (β)=0.01, Avg. K=2.612026.05 | 60.42 | — | |
| Qwen1.5-1.8B# Non-Emb Params=1.2B2024.07 | 57.8 | — | |
| Llama-3-SynEevaluation_mode=Few-shot2024.07 | 57.34 | — | |
| openPangu-Embedded RL# Total Params=1B, Training Stage=RL2026.05 | 56.53 | — | |
| openPangu-Embedded KD# Total Params=1B, Training Stage=KD (NPD)2026.05 | 55.43 | — | |
| Qwen2-0.5B# Non-Emb Params=0.3B2024.07 | 55.1 | — | |
| Qwen2.5# Total Params=1.5B2026.05 | 53.98 | — | |
| Mixtral-8x22B2024.07 | 53.4 | — | |
| openPangu-Embedded SFT# Total Params=1B, Training Stage=SFT2026.05 | 53.1 | — | |
| Llama-3-Chinese-8Bevaluation_mode=Few-shot2024.07 | 51.2 | — | |
| Llama-3-8Bevaluation_mode=Few-shot2024.07 | 51.03 | — | |
| DyMoEModel=Mixtral-8×7B, High / Low Bits=4 / 2, Retention ratio (r)=0.92026.03 | 50.53 | — | |
| DyMoEModel=Mixtral-8×7B, High / Low Bits=4 / 0, Retention ratio (r)=1.02026.03 | 50.44 | — | |
| DyMoEModel=Mixtral-8×7B, High / Low Bits=4 / 2, Retention ratio (r)=1.02026.03 | 50.44 | — | |
| DyMoEModel=Mixtral-8×7B, High / Low Bits=4 / 2, Retention ratio (r)=0.752026.03 | 49.47 | — | |
| DyMoEModel=Mixtral-8×7B, High / Low Bits=4 / 0, Retention ratio (r)=0.92026.03 | 49.4 | — | |
| DyMoEModel=Mixtral-8×7B, High / Low Bits=4 / 0, Retention ratio (r)=0.752026.03 | 48.22 | — | |
| DeepSeek 7B (Dense)# Shot=5-shot, # Total Params=6.9B, # Activated Params=6.9B, FLOPs per 4K Tokens=183.5T, # Training Tokens=2T2024.01 | 47.2 | — | |
| MAmmoTH2-8Bevaluation_mode=Few-shot2024.07 | 45.9 | — | |
| Mistral-7B-v0.3evaluation_mode=Few-shot2024.07 | 43.72 | — | |
| Qwen3# Total Params=0.6B2026.05 | 42.94 | — | |
| DeepSeekMoE 16B# Shot=5-shot, # Total Params=16.4B, # Activated Params=2.8B, FLOPs per 4K Tokens=74.4T, # Training Tokens=2T2024.01 | 42.5 | — | |
| DCLM-7Bevaluation_mode=Few-shot2024.07 | 40.89 | — | |
| DynMoEModel=Qwen1.5-MoE-A2.7B, Avg. K=30.062026.05 | 39.35 | — | |
| +49-class Div.Class Divisions=49-class, Training Tokens=150B, Model Size=15B-A1.5B2026.02 | 36.58 | — | |
| +3-class Div.Class Divisions=3-class, Training Tokens=150B, Model Size=15B-A1.5B2026.02 | 36.52 | — | |
| DynMoEModel=Qwen3-30B-A3B, Avg. K=61.662026.05 | 36.51 | — | |
| Baseline MoEClass Divisions=None, Training Tokens=150B, Model Size=15B-A1.5B2026.02 | 35.89 | — | |
| Gemma3# Total Params=1B2026.05 | 31.57 | — | |
| Galactica-6.7Bevaluation_mode=Few-shot2024.07 | 25.53 | — | |
| MiniMind-3Params=26M2026.05 | 25.3 | — | |
| MiniWinParams=26M2026.05 | 25 | — | |
| Phi-2# Non-Emb Params=2.5B2024.07 | 24.2 | — | |
| Llama3.2# Total Params=1B2026.05 | 10.81 | — | |
| DynMoEModel=DeepSeekV2-Lite, Avg. K=30.502026.05 | 0.41 | — | |
| Baichuan-7BModel Size=7B2023.08 | — | 44.4 | |
| Baichuan2-7BModel Size=7B2023.08 | — | 57.1 | |
| ChatGLM2-6BModel Size=6B2023.08 | — | 48.8 | |
| DeepSeek 67B (Dense)# Shot=5-shot2024.01 | — | 40.6 | |
| DeepSeek Chat 7B# Shot=0-shot, Total Params=6.9B, Activated Params=6.9B, FLOPs per 4K Tokens=183.5T2024.01 | — | 51.2 | |
| DeepSeekMoE 142B (Half Activated)# Shot=5-shot2024.01 | — | 31.9 | |
| DeepSeekMoE 145B# Shot=5-shot2024.01 | — | 35.9 | |
| DeepSeekMoE 16B# Shot=5-shot, # Total Params=16.4B, # Activated Params=2.8B, FLOPs per 4K Tokens=74.4T, # Training Tokens=2T2024.01 | — | 42.5 | |
| DeepSeekMoE Chat 16B# Shot=0-shot, Total Params=16.4B, Activated Params=2.8B, FLOPs per 4K Tokens=74.4T2024.01 | — | 49.3 | |
| Expert Divergence LearningModel Size=15B-A1.5B, Scheme=3-class2026.02 | — | 35.16 | |
| Expert Divergence LearningModel Size=15B-A1.5B, Scheme=49-class2026.02 | — | 35.17 | |
| Expert Divergence LearningModel Size=8B-A0.8B, Scheme=3-class2026.02 | — | 33.4 | |
| Expert Divergence LearningModel Size=8B-A0.8B, Scheme=49-class2026.02 | — | 34.5 | |
| Expert Divergence LearningModel Size=3B-A0.3B, Scheme=3-class2026.02 | — | 33.07 | |
| Expert Divergence LearningModel Size=3B-A0.3B, Scheme=49-class2026.02 | — | 32.91 | |
| GShard 137B# Shot=5-shot2024.01 | — | 25.4 | |
| InternLM-7BModel Size=7B2023.08 | — | 51.8 | |
| LLAMA-7BModel Size=7B2023.08 | — | 26.8 | |
| LLaMA2 7B# Shot=5-shot, # Total Params=6.7B, # Activated Params=6.7B, FLOPs per 4K Tokens=187.9T, # Training Tokens=2T2024.01 | — | 32.6 | |
| LLaMA2 SFT 7B# Shot=0-shot, Total Params=6.7B, Activated Params=6.7B, FLOPs per 4K Tokens=187.9T2024.01 | — | 36.9 | |
| LLAMA2-7BModel Size=7B2023.08 | — | 31.8 | |
| MoEModel Size=15B-A1.5B, Training Strategy=Baseline2026.02 | — | 34.64 | |
| MoEModel Size=8B-A0.8B, Training Strategy=Baseline2026.02 | — | 34.22 | |
| MoEModel Size=3B-A0.3B, Training Strategy=Baseline2026.02 | — | 32.73 | |
| Qwen-7BModel Size=7B, Status=final released2023.08 | — | 62.2 | |
| Qwen-VLInitialization=Intermediate Qwen-7B checkpoint2023.08 | — | 49.5 |