Summarization on CNN/DM (Completeness, Conciseness, Factualness)
5Completeness ScoreQuestion-Answer Guided (QAG)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Question-Answer Guided (QAG)Model Type=LLM (GPT-4.1)2025.12 | 5 | 4.07 | 5 | |
| Decomposition (Deco)Model Type=LLM (GPT-4.1)2025.12 | 4.98 | 4.21 | 5 | |
| Iterative Refine (IR)Model Type=LLM (GPT-4.1)2025.12 | 4.97 | 4.55 | 5 | |
| Chain-of-Thought (CoT)Model Type=LLM (GPT-4.1)2025.12 | 4.9 | 4.23 | 4.98 | |
| Self-Consistency (SC)Model Type=LLM (GPT-4.1)2025.12 | 4.89 | 4.63 | 5 | |
| Extract-to-Abstract (E2A)Model Type=LLM (GPT-4.1)2025.12 | 4.89 | 4.88 | 5 | |
| Plan-then-Write (Plan)Model Type=LLM (GPT-4.1)2025.12 | 4.88 | 4.14 | 5 | |
| VanillaModel Type=LLM (GPT-4.1)2025.12 | 4.82 | 4.6 | 4.97 | |
| GPT-5Model Type=Large Reasoning Models (LRM)2025.12 | 4.79 | 4.47 | 5 | |
| Cited Summarization (Cite)Model Type=LLM (GPT-4.1)2025.12 | 4.77 | 4.89 | 5 | |
| o1Model Type=Large Reasoning Models (LRM)2025.12 | 4.76 | 4.77 | 4.93 | |
| o3Model Type=Large Reasoning Models (LRM)2025.12 | 4.7 | 4.74 | 4.93 |