Abstractive Summarization on BookSum sampled (test)
17.71ROUGE ScoreIterative Refine (IR)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Iterative Refine (IR)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 17.71 | 85.84 | |
| Self-Consistency (SC)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 17.68 | 85.76 | |
| o3Shot=0-shot, Model Category=Large Reasoning Models (LRM)2025.12 | 17.68 | 85.18 | |
| VanillaShot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 17.49 | 85.76 | |
| Question-Answer Guided (QAG)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 17.49 | 85.47 | |
| Extract-to-Abstract (E2A)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 17.49 | 85.77 | |
| GPT-5Shot=0-shot, Model Category=Large Reasoning Models (LRM)2025.12 | 17.39 | 84.75 | |
| Decomposition (Deco)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 17.15 | 85.66 | |
| o1Shot=0-shot, Model Category=Large Reasoning Models (LRM)2025.12 | 17.12 | 84.69 | |
| Cited Summarization (Cite)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 16.34 | 85.64 | |
| Chain-of-Thought (CoT)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 14.87 | 84.65 | |
| Plan-then-Write (Plan)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 14.4 | 84.69 |