Language Understanding on C-Eval
87.7C-Eval ScoreQwen2-57B-A14B
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2-57B-A14BArchitecture=MoE, # Act Params=14B, # Params=57B2024.07 | 87.7 | |
| Qwen1.5-32BArchitecture=Dense, # Act Params=34B, # Params=34B2024.07 | 83.5 | |
| Qwen2-7BParameters=7B2024.07 | 83.2 | |
| Qwen1.5-7BParameters=7B2024.07 | 74.1 | |
| KEELArchitecture=512 Layers / 3B Params, Peak Learning Rate=4.5 x 10^-3, Pre-training Tokens=1T, Evaluation Protocol (Shots)=5-shot2026.01 | 69.5 | |
| CPTSize=70B, Evaluation Mode=5-shot2024.09 | 68.36 | |
| Llama-3Size=70B, Evaluation Mode=5-shot2024.09 | 67.75 | |
| Pre-LNArchitecture=512 Layers / 3B Params, Peak Learning Rate=3.0 x 10^-3, Pre-training Tokens=1T, Evaluation Protocol (Shots)=5-shot2026.01 | 66.2 | |
| Qwen-7BModel Size=7B, Status=final released2023.08 | 63.5 | |
| Llama-3-SynEevaluation_mode=Few-shot2024.07 | 58.24 | |
| Baichuan2-7BModel Size=7B2023.08 | 54 | |
| InternLM-7BModel Size=7B2023.08 | 52.8 | |
| ChatGLM2-6BModel Size=6B2023.08 | 51.7 | |
| CPTSize=8B, Evaluation Mode=5-shot2024.09 | 51.12 | |
| Qwen-VLInitialization=Intermediate Qwen-7B checkpoint2023.08 | 51.1 | |
| Llama-3-ChineseSize=8B, Evaluation Mode=5-shot2024.09 | 50.5 | |
| Llama-3-Chinese-8Bevaluation_mode=Few-shot2024.07 | 50.14 | |
| Llama-3-8BParameters=8B2024.07 | 49.5 | |
| Llama-3-8Bevaluation_mode=Few-shot2024.07 | 49.43 | |
| Llama-3Size=8B, Evaluation Mode=5-shot2024.09 | 49.38 | |
| Qwen-7BModel Size=7B, Status=intermediate, Usage=LLM initialization for Qwen-VL2023.08 | 48.5 | |
| Mistral-7BParameters=7B2024.07 | 47.4 | |
| MAmmoTH2-8Bevaluation_mode=Few-shot2024.07 | 46.56 | |
| Gemma-7BParameters=7B2024.07 | 43.6 | |
| Baichuan-7BModel Size=7B2023.08 | 42.8 | |
| Mistral-7B-v0.3evaluation_mode=Few-shot2024.07 | 42.74 | |
| DCLM-7Bevaluation_mode=Few-shot2024.07 | 41.24 | |
| LLAMA2-7BModel Size=7B2023.08 | 32.5 | |
| Galactica-6.7Bevaluation_mode=Few-shot2024.07 | 26.72 |