Large Language Model Capability Evaluation on Nine-capability breakdown C1-C9
27.49Capability 1 ScoreGLM-5.1
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| GLM-5.12026.05 | 27.49 | 48.48 | 20.05 | 84.76 | 99.47 | 69.58 | 68.4 | 42.34 | 66.06 | |
| Claude-Sonnet-4.62026.05 | 26.6 | 44.19 | 19.77 | 88.81 | 99.54 | 66.67 | 68.49 | 44.59 | 70.08 | |
| GPT-5.3-Codex2026.05 | 25.32 | 42.9 | 19.89 | 80.51 | 98.54 | 55.94 | 77.12 | 42.34 | 68.09 | |
| MiniMax-M2.72026.05 | 24.89 | 49.51 | 19.69 | 86.17 | 99.82 | 58.96 | 68.72 | 43.9 | 60.77 | |
| Qwen-3.5-122B-A10B2026.05 | 22.46 | 53.05 | 19.38 | 88.83 | 99.52 | 60.42 | 68.81 | 42.62 | 55.72 | |
| Gemini-3-Flash-Prev.2026.05 | 22.18 | 52.17 | 19.34 | 82.85 | 94.08 | 66.15 | 64.51 | 42.06 | 55.14 |