Abstractive Summarization on Average across 8 datasets
19.89ROUGE ScoreSelf-Consistency (SC)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Self-Consistency (SC)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 19.89 | 86.17 | |
| Iterative Refine (IR)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 19.77 | 85.7 | |
| Extract-to-Abstract (E2A)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 19.69 | 86.07 | |
| VanillaShot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 19.24 | 85.17 | |
| Cited Summarization (Cite)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 18.97 | 86.07 | |
| Question-Answer Guided (QAG)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 18.68 | 85.18 | |
| o1Shot=0-shot, Model Category=Large Reasoning Models (LRM)2025.12 | 17.98 | 83.77 | |
| Chain-of-Thought (CoT)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 17.47 | 84.92 | |
| Decomposition (Deco)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 17.43 | 85.13 | |
| GPT-5Shot=0-shot, Model Category=Large Reasoning Models (LRM)2025.12 | 16.99 | 82.51 | |
| Plan-then-Write (Plan)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 16.83 | 84.94 | |
| o3Shot=0-shot, Model Category=Large Reasoning Models (LRM)2025.12 | 16.54 | 83.77 |