Abstractive Summarization on WikiHow sampled (test)
17.58ROUGECited Summarization (Cite)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Cited Summarization (Cite)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 17.58 | 85.3 | |
| Extract-to-Abstract (E2A)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 17.21 | 84.85 | |
| VanillaShot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 16.75 | 83.85 | |
| Self-Consistency (SC)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 16.75 | 84.89 | |
| Iterative Refine (IR)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 16.22 | 84.67 | |
| o1Shot=0-shot, Model Category=Large Reasoning Models (LRM)2025.12 | 16.04 | 82.82 | |
| o3Shot=0-shot, Model Category=Large Reasoning Models (LRM)2025.12 | 14.87 | 83.96 | |
| Question-Answer Guided (QAG)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 14.83 | 84.12 | |
| Chain-of-Thought (CoT)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 14.72 | 84.22 | |
| Decomposition (Deco)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 14.51 | 84.4 | |
| Plan-then-Write (Plan)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 14.43 | 84.15 | |
| GPT-5Shot=0-shot, Model Category=Large Reasoning Models (LRM)2025.12 | 13.75 | 79.46 |