Abstractive Summarization on Multi-News 56k samples (test)
20.72ROUGE ScoreQuestion-Answer Guided (QAG)
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Question-Answer Guided (QAG)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 20.72 | — | — | — | 85.44 | |
| Iterative Refine (IR)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 20.67 | — | — | — | 85.64 | |
| Chain-of-Thought (CoT)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 19.95 | — | — | — | 85.44 | |
| VanillaShot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 19.79 | — | — | — | 85.41 | |
| Self-Consistency (SC)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 19.47 | — | — | — | 85.5 | |
| Extract-to-Abstract (E2A)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 19.05 | — | — | — | 85.43 | |
| GPT-5Shot=0-shot, Model Category=Large Reasoning Models (LRM)2025.12 | 18.81 | — | — | — | 83.67 | |
| Plan-then-Write (Plan)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 18.33 | — | — | — | 85.09 | |
| Decomposition (Deco)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 18.26 | — | — | — | 85.09 | |
| o3Shot=0-shot, Model Category=Large Reasoning Models (LRM)2025.12 | 17.99 | — | — | — | 84.67 | |
| o1Shot=0-shot, Model Category=Large Reasoning Models (LRM)2025.12 | 17.97 | — | — | — | 84.17 | |
| Cited Summarization (Cite)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 17.89 | — | — | — | 85.26 | |
| PEGASUSBASEModel Size=Base2019.12 | — | 42.24 | 13.27 | 21.44 | — | |
| PEGASUSLARGEModel Size=Large, Pre-training Corpus=C42019.12 | — | 46.74 | 17.95 | 24.26 | — | |
| PEGASUSLARGEModel Size=Large, Pre-training Corpus=HugeNews2019.12 | — | 47.52 | 18.72 | 24.91 | — | |
| Previous SOTA2019.12 | — | 43.47 | 14.89 | 17.41 | — | |
| TransformerBASEModel Size=Base2019.12 | — | 34.36 | 5.42 | 15.75 | — |