Text Summarization on CNN/DailyMail LLM Reference (test)
45.1ROUGE-1Mistral-Large
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| Mistral-Large2026.04 | 45.1 | 18.8 | 29.4 | 55.9 | 71.2 | -2.48 | |
| LLMs Average2026.04 | 42 | 17.3 | 27.2 | 54.5 | 69.7 | -2.59 | |
| Gemini-1.5-pro2026.04 | 41.5 | 16.7 | 27.2 | 53.9 | 69.7 | -2.67 | |
| GPT-4o2026.04 | 41.2 | 15.9 | 26 | 54.7 | 69.4 | -2.78 |