Summarization on CNN/DM (test) (ROUGE, BS)
23.39ROUGEVanilla
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| VanillaBackbone=GPT-4.1, Shot-setting=2-shot2025.12 | 23.39 | 83.82 | |
| Iterative Refine (IR)Backbone=GPT-4.1, Shot-setting=2-shot2025.12 | 22.8 | 86.27 | |
| Self-Consistency (SC)Backbone=GPT-4.1, Shot-setting=2-shot2025.12 | 21.96 | 84.58 | |
| Chain-of-Thought (CoT)Backbone=GPT-4.1, Shot-setting=2-shot2025.12 | 21.68 | 84.32 | |
| Extract-to-Abstract (E2A)Backbone=GPT-4.1, Shot-setting=2-shot2025.12 | 21.66 | 85.19 | |
| Cited Summarization (Cite)Backbone=GPT-4.1, Shot-setting=2-shot2025.12 | 21.58 | 84.66 | |
| o1Backbone=Large Reasoning Model, Shot-setting=2-shot2025.12 | 21.54 | 83.56 | |
| Question-Answer Guided (QAG)Backbone=GPT-4.1, Shot-setting=2-shot2025.12 | 20.91 | 84.09 | |
| Decomposition (Deco)Backbone=GPT-4.1, Shot-setting=2-shot2025.12 | 19.96 | 84.96 | |
| Plan-then-Write (Plan)Backbone=GPT-4.1, Shot-setting=2-shot2025.12 | 19.89 | 85.75 | |
| GPT-5Backbone=Large Reasoning Model, Shot-setting=2-shot2025.12 | 19.8 | 82.41 | |
| o3Backbone=Large Reasoning Model, Shot-setting=2-shot2025.12 | 18.28 | 82.47 |