Summarization on BookSum (test)
5Comp ScoreChain-of-Thought (CoT)
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Chain-of-Thought (CoT)Base Model=GPT-4.12025.12 | 5 | 3.17 | 4.98 | — | — | |
| Decomposition (Deco)Base Model=GPT-4.12025.12 | 5 | 4.23 | 5 | — | — | |
| Iterative Refine (IR)Base Model=GPT-4.12025.12 | 5 | 4.53 | 5 | — | — | |
| COTModel Category=LLM (GPT-4.1)2025.12 | 5 | 3.17 | 4.98 | — | — | |
| DecoModel Category=LLM (GPT-4.1)2025.12 | 5 | 4.23 | 5 | — | — | |
| IRModel Category=LLM (GPT-4.1)2025.12 | 5 | 4.53 | 5 | — | — | |
| Self-Consistency (SC)Base Model=GPT-4.12025.12 | 4.98 | 4.63 | 4.97 | — | — | |
| SCModel Category=LLM (GPT-4.1)2025.12 | 4.98 | 4.63 | 4.97 | — | — | |
| o3Model Category=Large Reasoning Models (LRM)2025.12 | 4.97 | 4.73 | 5 | — | — | |
| o3Model Category=Large Reasoning Models (LRM)2025.12 | 4.97 | 4.73 | 5 | — | — | |
| Plan-then-Write (Plan)Base Model=GPT-4.12025.12 | 4.95 | 4.04 | 4.99 | — | — | |
| PlanModel Category=LLM (GPT-4.1)2025.12 | 4.95 | 4.04 | 4.99 | — | — | |
| o1Model Category=Large Reasoning Models (LRM)2025.12 | 4.94 | 4.87 | 4.98 | — | — | |
| o1Model Category=Large Reasoning Models (LRM)2025.12 | 4.94 | 4.87 | 4.98 | — | — | |
| Cited Summarization (Cite)Base Model=GPT-4.12025.12 | 4.93 | 4.88 | 4.97 | — | — | |
| GPT-5Model Category=Large Reasoning Models (LRM)2025.12 | 4.93 | 4.25 | 4.99 | — | — | |
| CiteModel Category=LLM (GPT-4.1)2025.12 | 4.93 | 4.88 | 4.97 | — | — | |
| GPT-5Model Category=Large Reasoning Models (LRM)2025.12 | 4.93 | 4.25 | 4.99 | — | — | |
| VanillaBase Model=GPT-4.12025.12 | 4.92 | 4.47 | 4.94 | — | — | |
| VanillaModel Category=LLM (GPT-4.1)2025.12 | 4.92 | 4.47 | 4.94 | — | — | |
| Extract-to-Abstract (E2A)Base Model=GPT-4.12025.12 | 4.24 | 4.83 | 4.95 | — | — | |
| E2AModel Category=LLM (GPT-4.1)2025.12 | 4.24 | 4.83 | 4.95 | — | — | |
| Question-Answer Guided (QAG)Base Model=GPT-4.12025.12 | 2.03 | 2.5 | 4.85 | — | — | |
| QA-GModel Category=LLM (GPT-4.1)2025.12 | 2.03 | 2.5 | 4.85 | — | — | |
| Chain-of-Thought (CoT)Backbone=GPT-4.1, Shot-setting=2-shot2025.12 | — | — | — | 19.24 | 86.14 | |
| Cited Summarization (Cite)Backbone=GPT-4.1, Shot-setting=2-shot2025.12 | — | — | — | 17.01 | 84.09 | |
| Decomposition (Deco)Backbone=GPT-4.1, Shot-setting=2-shot2025.12 | — | — | — | 17.8 | 85.83 | |
| Extract-to-Abstract (E2A)Backbone=GPT-4.1, Shot-setting=2-shot2025.12 | — | — | — | 18.34 | 85.92 | |
| GPT-5Backbone=Large Reasoning Model, Shot-setting=2-shot2025.12 | — | — | — | 18.23 | 84.05 | |
| Iterative Refine (IR)Backbone=GPT-4.1, Shot-setting=2-shot2025.12 | — | — | — | 19.32 | 86.11 | |
| o1Backbone=Large Reasoning Model, Shot-setting=2-shot2025.12 | — | — | — | 16.85 | 84.66 | |
| o3Backbone=Large Reasoning Model, Shot-setting=2-shot2025.12 | — | — | — | 17.36 | 82.51 | |
| Plan-then-Write (Plan)Backbone=GPT-4.1, Shot-setting=2-shot2025.12 | — | — | — | 15.71 | 85.02 | |
| Question-Answer Guided (QAG)Backbone=GPT-4.1, Shot-setting=2-shot2025.12 | — | — | — | 18.31 | 85.7 | |
| Self-Consistency (SC)Backbone=GPT-4.1, Shot-setting=2-shot2025.12 | — | — | — | 19 | 85.99 | |
| VanillaBackbone=GPT-4.1, Shot-setting=2-shot2025.12 | — | — | — | 19.33 | 85.32 |