Text Quality Meta-evaluation on SummEval & Topical-Chat Combined
69.5Overall ScoreDeepSeek-V3
Evaluation Results
| Method | Links | |
|---|---|---|
| DeepSeek-V32025.02 | 69.5 | |
| GPT-4o2025.02 | 69.4 | |
| GPT-4 Turbo2025.02 | 68.9 | |
| GPT-4o mini2025.02 | 68.4 | |
| Qwen-2.5-72B2025.02 | 67.4 | |
| Gemma-2-27B2025.02 | 66.9 | |
| CompassJudger-32B2025.02 | 66.7 | |
| Phi-4-14B2025.02 | 65.8 | |
| Llama-3.1-70B2025.02 | 65 | |
| GPT-3.5 Turbo2025.02 | 62.5 | |
| Prometheus-2-8x7B2025.02 | 59.7 | |
| CRITIQUELLM-6B2025.02 | 59.6 | |
| Prometheus-2-7B2025.02 | 59 | |
| Auto-J-13B2025.02 | 51.4 | |
| Prometheus-13B2025.02 | 48.4 | |
| Themis-8B2025.02 | 41.7 |