Chinese Language Understanding on C-Eval
92.5AccuracyKimi-K2
Evaluation Results
| Method | Links | |
|---|---|---|
| Kimi-K2Model Variant=Base, # Shots=5-shot, # Activated Params=32B, # Total Params=1043B2026.02 | 92.5 | |
| Qwen2-72B2024.07 | 91 | |
| DeepSeek V3.2Model Variant=Exp Base, # Shots=5-shot, # Activated Params=37B, # Total Params=671B2026.02 | 91 | |
| DeepSeek V3.1Model Variant=Base, # Shots=5-shot, # Activated Params=37B, # Total Params=671B2026.02 | 90 | |
| Step 3.5 FlashModel Variant=Base, # Shots=5-shot, # Activated Params=11B, # Total Params=196B2026.02 | 89.6 | |
| Qwen1.5-110B2024.07 | 89.1 | |
| MiMo-V2 FlashModel Variant=Base, # Shots=5-shot, # Activated Params=15B, # Total Params=309B2026.02 | 87.9 | |
| GLM-4.5Model Variant=Base, # Shots=5-shot, # Activated Params=32B, # Total Params=355B2026.02 | 86.9 | |
| Qwen1.5-72B2024.07 | 84.1 | |
| Qwen2.5 (base)Params=7B, Tokens=18T, Complexity Type=Quadratic, Evaluation Framework=HuggingFace, Evaluation Protocol=perplexity-based2025.09 | 81.6 | |
| Qwen3-4BParams=4B2025.12 | 78.5 | |
| Qwen2.5-3BParams=3B2025.12 | 74.65 | |
| SDAR 8BParameters=8B, Sampling=Standard, Note=Non-official replication2025.12 | 72.7 | |
| NBDiff-7B-INSTRUCTParameters=7B, Training Protocol=Instruct, Sampling=Standard2025.12 | 72.5 | |
| Qwen2-1.5BParams=1.5B2025.12 | 71.29 | |
| Qwen2-1.5B# Non-Emb Params=1.2B2024.07 | 70.6 | |
| SpikingBrain-7BParams=7B, Tokens=+150B, Complexity Type=Linear, Evaluation Framework=HuggingFace, Evaluation Protocol=perplexity-based2025.09 | 69.8 | |
| Qwen2.5-1.5BParams=1.5B2025.12 | 68.63 | |
| Qwen3-1.7BParams=1.7B2025.12 | 66.7 | |
| LLaDA2.0-mini preview 16BA1BParameters=16BA1B, Sampling=Standard2025.12 | 66.5 | |
| HSA-ULTraining Strategy=Annealing, Architecture=MoE, Total Params=8B, Activated Params=1B, Training Tokens=8T2025.11 | 65.98 | |
| CAMELTraining Objective=Balanced, Sampling Strategy=Hourglass2026.03 | 65.3 | |
| Llama-3-70B2024.07 | 65.2 | |
| HySparse# Shots=5-shot, Model Architecture=80B MoE (Hybrid 1:11), Attention Variant=HySparse2026.02 | 65 | |
| Full-Attn# Shots=5-shot, Model Architecture=80B MoE (Hybrid 1:11), Attention Variant=Full-Attn2026.02 | 64.6 | |
| LLaDA-MoE 7B-A1BParameters=7B-A1B, Sampling=Standard2025.12 | 63.9 | |
| Human DesignedTraining Objective=Balanced, Sampling Strategy=Rectangle2026.03 | 63.7 | |
| OPTIMERBase Model=Gemma 3 27B2026.03 | 63.52 | |
| DMLTraining Objective=Balanced, Sampling Strategy=Rectangle2026.03 | 63.1 | |
| SODMTraining Objective=Balanced, Sampling Strategy=Rectangle2026.03 | 62.8 | |
| Model-size agnosticTraining Objective=Balanced, Sampling Strategy=Rectangle2026.03 | 62.7 | |
| OPTIMERBase Model=SEA-LION v42026.03 | 62.56 | |
| Qwen1.5-1.8B# Non-Emb Params=1.2B2024.07 | 59.7 | |
| Hybrid SWA# Shots=5-shot, Model Architecture=80B MoE (Hybrid 1:11), Attention Variant=Hybrid SWA2026.02 | 58.8 | |
| HSA-ULTraining Strategy=Base, Architecture=MoE, Total Params=8B, Activated Params=1B, Training Tokens=8T2025.11 | 58.36 | |
| Qwen2-0.5B# Non-Emb Params=0.3B2024.07 | 58.2 | |
| Dream-v0 Instruct-7BParameters=7B, Training Protocol=Instruct, Sampling=Standard2025.12 | 58 | |
| Qwen3-0.6BParams=0.6B2025.12 | 57.03 | |
| TRM-MoETraining Strategy=Base, Architecture=MoE, Total Params=8B, Activated Params=1B, Training Tokens=8T2025.11 | 56.87 | |
| Mixtral-8x22B2024.07 | 54.6 | |
| Qwen3Training Strategy=Annealing, Architecture=Dense, Total Params=0.6B, Activated Params=0.6B, Training Tokens=36T2025.11 | 54.57 | |
| Qwen2.5Training Strategy=Annealing, Architecture=Dense, Total Params=0.5B, Activated Params=0.5B, Training Tokens=18T2025.11 | 54.17 | |
| YuLan-Mini-2.4BParams=2.4B2025.12 | 52.32 | |
| HySparse# Shots=5-shot, Model Architecture=7B Dense (Hybrid 1:3), Attention Variant=HySparse2026.02 | 52.2 | |
| Llama3.1Params=8B, Tokens=15T, Complexity Type=Quadratic, Evaluation Framework=HuggingFace, Evaluation Protocol=perplexity-based2025.09 | 51.46 | |
| SEA-LION v4 27B IT2026.03 | 50.89 | |
| SmolLM3-3BParams=3B2025.12 | 50.84 | |
| Full-Attn# Shots=5-shot, Model Architecture=7B Dense (Hybrid 1:3), Attention Variant=Full-Attn2026.02 | 50.6 | |
| Hybrid SWA# Shots=5-shot, Model Architecture=7B Dense (Hybrid 1:3), Attention Variant=Hybrid SWA2026.02 | 50.6 | |
| MistralParams=7B, Complexity Type=Linear, Evaluation Framework=HuggingFace, Evaluation Protocol=perplexity-based2025.09 | 47.04 | |
| PCMind-2.1-Kaiyuan-2BParams=2B2025.12 | 46.3 | |
| llama-3.2-3BParams=3B2025.12 | 45.67 | |
| HSA-ULTraining Strategy=Annealing, Architecture=Dense, Total Params=0.5B, Activated Params=0.5B, Training Tokens=4T2025.11 | 44.3 | |
| Falcon-MambaParams=7B, Tokens=5.8T, Complexity Type=Linear, Evaluation Framework=HuggingFace, Evaluation Protocol=perplexity-based2025.09 | 41.93 | |
| gemma2-2BParams=2B2025.12 | 41.35 | |
| NITPModel scale=9bA1b, Evaluation=few-shot, Context length=81922026.05 | 40.72 | |
| NTPModel scale=9bA1b, Evaluation=few-shot, Context length=81922026.05 | 38.97 | |
| Gemma 3 27B IT2026.03 | 37.96 | |
| Zamba-v1Params=7B, Tokens=1T, Complexity Type=Hybrid, Evaluation Framework=HuggingFace, Evaluation Protocol=perplexity-based2025.09 | 36.4 | |
| SmolLM2-1.7BParams=1.7B2025.12 | 35.06 | |
| NTPModel scale=3bA0.5b, Evaluation=few-shot, Context length=81922026.05 | 33.53 | |
| NITPModel scale=1.9bA0.3b, Evaluation=few-shot, Context length=81922026.05 | 33.16 | |
| NITPModel scale=3bA0.5b, Evaluation=few-shot, Context length=81922026.05 | 33.12 | |
| NTPModel scale=1.9bA0.3b, Evaluation=few-shot, Context length=81922026.05 | 32.31 | |
| OLMo-2-0425-1BParams=1B2025.12 | 30.53 | |
| llama-3.2-1BParams=1B2025.12 | 29.82 | |
| Gemma-2B# Non-Emb Params=2.0B2024.07 | 28 | |
| Phi-2# Non-Emb Params=2.5B2024.07 | 23.4 |