Language Modeling on General Benchmarks (C-Eval, IFEval, MATH-500, LCB, H-Swag, SST-5, CrossNER) (test)
80.2C-Eval ScoreQwen2.5-14B-Ins.
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Qwen2.5-14B-Ins.Model Group=Base Models2026.06 | 80.2 | 79.6 | 80.6 | 20.4 | 85.9 | 54.3 | 59.7 | 65.8 | |
| YouZhi-14BModel Group=Our Models2026.06 | 78.9 | 79.5 | 72.4 | 25.5 | 92.8 | 65.2 | 61.5 | 68 | |
| DianJin-R1-7BModel Group=Baselines2026.06 | 74.5 | 48.8 | 76.8 | 9.3 | 24.4 | 42.9 | 34.2 | 44.4 | |
| OpenPangu-7BModel Group=Base Models2026.06 | 73.4 | 77.8 | 76.6 | 34.6 | 67.8 | 51.5 | 55.2 | 62.4 | |
| YouZhi-7BModel Group=Our Models2026.06 | 72.9 | 77.1 | 76 | 30.1 | 90.5 | 62.6 | 58.1 | 66.8 | |
| YiZhao-12B-ChatModel Group=Baselines2026.06 | 70.3 | 43.9 | 46.5 | 3.6 | 70.2 | 54.3 | 49.8 | 48.4 |