Summarization on MultiNews (test)
4.98ComprehensivenessChain-of-Thought (CoT)
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| Chain-of-Thought (CoT)Base Model=GPT-4.12025.12 | 4.98 | — | — | — | — | 4.24 | 5 | |
| COTModel Category=LLM (GPT-4.1)2025.12 | 4.98 | — | — | — | — | 4.24 | 5 | |
| Iterative Refine (IR)Base Model=GPT-4.12025.12 | 4.96 | — | — | — | — | 4.51 | 4.97 | |
| IRModel Category=LLM (GPT-4.1)2025.12 | 4.96 | — | — | — | — | 4.51 | 4.97 | |
| Self-Consistency (SC)Base Model=GPT-4.12025.12 | 4.95 | — | — | — | — | 4.59 | 4.98 | |
| SCModel Category=LLM (GPT-4.1)2025.12 | 4.95 | — | — | — | — | 4.59 | 4.98 | |
| VanillaBase Model=GPT-4.12025.12 | 4.94 | — | — | — | — | 4.56 | 4.99 | |
| VanillaModel Category=LLM (GPT-4.1)2025.12 | 4.94 | — | — | — | — | 4.56 | 4.99 | |
| Decomposition (Deco)Base Model=GPT-4.12025.12 | 4.92 | — | — | — | — | 4.1 | 4.98 | |
| DecoModel Category=LLM (GPT-4.1)2025.12 | 4.92 | — | — | — | — | 4.1 | 4.98 | |
| Plan-then-Write (Plan)Base Model=GPT-4.12025.12 | 4.91 | — | — | — | — | 4.14 | 4.95 | |
| PlanModel Category=LLM (GPT-4.1)2025.12 | 4.91 | — | — | — | — | 4.14 | 4.95 | |
| o3Model Category=Large Reasoning Models (LRM)2025.12 | 4.86 | — | — | — | — | 4.67 | 4.97 | |
| o3Model Category=Large Reasoning Models (LRM)2025.12 | 4.86 | — | — | — | — | 4.67 | 4.97 | |
| Cited Summarization (Cite)Base Model=GPT-4.12025.12 | 4.83 | — | — | — | — | 4.84 | 4.98 | |
| CiteModel Category=LLM (GPT-4.1)2025.12 | 4.83 | — | — | — | — | 4.84 | 4.98 | |
| o1Model Category=Large Reasoning Models (LRM)2025.12 | 4.8 | — | — | — | — | 4.63 | 4.92 | |
| o1Model Category=Large Reasoning Models (LRM)2025.12 | 4.8 | — | — | — | — | 4.63 | 4.92 | |
| GPT-5Model Category=Large Reasoning Models (LRM)2025.12 | 4.79 | — | — | — | — | 4.2 | 4.98 | |
| GPT-5Model Category=Large Reasoning Models (LRM)2025.12 | 4.79 | — | — | — | — | 4.2 | 4.98 | |
| Extract-to-Abstract (E2A)Base Model=GPT-4.12025.12 | 4.68 | — | — | — | — | 4.82 | 4.99 | |
| E2AModel Category=LLM (GPT-4.1)2025.12 | 4.68 | — | — | — | — | 4.82 | 4.99 | |
| Question-Answer Guided (QAG)Base Model=GPT-4.12025.12 | 1.93 | — | — | — | — | 3.97 | 4.98 | |
| QA-GModel Category=LLM (GPT-4.1)2025.12 | 1.93 | — | — | — | — | 3.97 | 4.98 | |
| BART-PTShots=100-shot2022.11 | — | 4.17 | 3.8 | 3.27 | 3.63 | — | — | |
| Goldtype=human reference2022.11 | — | 4.7 | 4.7 | 4.93 | 4.77 | — | — | |
| GPT-3.5Model=text-davinci-002, Shots=1-shot2022.11 | — | 4.97 | 4.73 | 3.07 | 2.73 | — | — | |
| PEGASUSShots=100-shot2022.11 | — | 4.23 | 3.95 | 3.53 | 3.72 | — | — | |
| UNISUMMShots=100-shot2022.11 | — | 4.63 | 4.17 | 4.07 | 4.3 | — | — |