Abstractive Summarization on CNN/DM sampled (test)
22.86ROUGE ScoreExtract-to-Abstract (E2A)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Extract-to-Abstract (E2A)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 22.86 | 87.01 | |
| Cited Summarization (Cite)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 22.76 | 87.16 | |
| Iterative Refine (IR)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 22.68 | 86.16 | |
| Self-Consistency (SC)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 22.23 | 86.94 | |
| VanillaShot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 21.27 | 84.24 | |
| o1Shot=0-shot, Model Category=Large Reasoning Models (LRM)2025.12 | 21.12 | 85.01 | |
| Chain-of-Thought (CoT)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 20.88 | 84.88 | |
| Question-Answer Guided (QAG)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 20.03 | 84.6 | |
| GPT-5Shot=0-shot, Model Category=Large Reasoning Models (LRM)2025.12 | 19.64 | 82.41 | |
| Decomposition (Deco)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 19.47 | 86.47 | |
| Plan-then-Write (Plan)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 18.94 | 84.7 | |
| o3Shot=0-shot, Model Category=Large Reasoning Models (LRM)2025.12 | 18.54 | 82.54 |