Abstractive Summarization on SAMSum sampled (test)
26.88ROUGE ScoreVanilla
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| VanillaShot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 26.88 | 88.85 | |
| Self-Consistency (SC)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 26.67 | 90.01 | |
| Extract-to-Abstract (E2A)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 26.51 | 89.85 | |
| Iterative Refine (IR)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 25.1 | 88.79 | |
| Question-Answer Guided (QAG)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 24.78 | 89.3 | |
| Cited Summarization (Cite)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 24.25 | 89.51 | |
| GPT-5Shot=0-shot, Model Category=Large Reasoning Models (LRM)2025.12 | 23.52 | 84.6 | |
| o1Shot=0-shot, Model Category=Large Reasoning Models (LRM)2025.12 | 23.5 | 87.59 | |
| o3Shot=0-shot, Model Category=Large Reasoning Models (LRM)2025.12 | 22.02 | 87.44 | |
| Decomposition (Deco)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 21.62 | 86.85 | |
| Chain-of-Thought (CoT)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 21.03 | 88.11 | |
| Plan-then-Write (Plan)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 19.57 | 88.39 |