Summarization on SAMSum (Completeness, Conciseness, Factualness)
4.98CompletenessPlan-then-Write (Plan)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Plan-then-Write (Plan)Model Type=LLM (GPT-4.1)2025.12 | 4.98 | 3.54 | 4.99 | |
| VanillaModel Type=LLM (GPT-4.1)2025.12 | 4.96 | 4.69 | 4.99 | |
| Question-Answer Guided (QAG)Model Type=LLM (GPT-4.1)2025.12 | 4.96 | 3.77 | 5 | |
| Iterative Refine (IR)Model Type=LLM (GPT-4.1)2025.12 | 4.96 | 4.46 | 5 | |
| Extract-to-Abstract (E2A)Model Type=LLM (GPT-4.1)2025.12 | 4.96 | 4.85 | 5 | |
| Decomposition (Deco)Model Type=LLM (GPT-4.1)2025.12 | 4.94 | 3.85 | 4.98 | |
| Cited Summarization (Cite)Model Type=LLM (GPT-4.1)2025.12 | 4.93 | 4.44 | 4.98 | |
| Self-Consistency (SC)Model Type=LLM (GPT-4.1)2025.12 | 4.92 | 4.73 | 5 | |
| Chain-of-Thought (CoT)Model Type=LLM (GPT-4.1)2025.12 | 4.91 | 4.31 | 4.97 | |
| o3Model Type=Large Reasoning Models (LRM)2025.12 | 4.88 | 4.81 | 4.96 | |
| o1Model Type=Large Reasoning Models (LRM)2025.12 | 4.82 | 4.61 | 4.94 | |
| GPT-5Model Type=Large Reasoning Models (LRM)2025.12 | 4.77 | 4.67 | 4.94 |