Multi-document Summarization on Multi-News (test) using ROUGE
21.7ROUGE-2UL20B
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| UL20Bmode=supervised finetuning2022.05 | 21.7 | — | — | — | — | — | |
| Z-Code++_LARGENumber of Parameters=710M2022.08 | 21.6 | — | — | — | — | — | |
| PRIMER2021.12 | 21.1 | 49.9 | — | 25.9 | — | — | |
| Xiao et al.status=previous state-of-the-art2022.05 | 21.1 | — | — | — | — | — | |
| PRIMERA*Evaluation Protocol=fully-supervised, Reference=Xiao et al. (2022)2022.08 | 21.1 | 49.9 | — | 25.9 | — | — | |
| PRIMER2022.08 | 21.1 | — | — | — | — | — | |
| PRIMERAtraining mode=fully supervised2021.10 | 21.05 | 49.94 | — | 25.85 | — | — | |
| PRIMERA2023.08 | 21.05 | 49.94 | — | 25.85 | — | — | |
| PRIMERAEvaluation Protocol=fully-supervised, Implementation=reproduced2022.08 | 20.6 | 50 | — | 25.5 | — | — | |
| SimCASBackbone=BART-large2023.08 | 20.47 | 49.4 | — | 25.96 | 65.4 | — | |
| CentrumEvaluation Protocol=fully-supervised2022.08 | 20.4 | 49 | — | 25.4 | — | — | |
| SimCASBackbone=BART-base2023.08 | 20.01 | 48.88 | — | 25.32 | 65.05 | — | |
| BART-Longlength limit=10002021.10 | 19.5 | 49.15 | — | 24.47 | — | — | |
| LongT5Model size=xl, Input length=8k, Attention mechanism=TGlobal2021.12 | 19.43 | 48.17 | — | 24.94 | — | — | |
| LongT5_XLARGENumber of Parameters=3B2022.08 | 19.4 | — | — | — | — | — | |
| BART-Long-Graph2021.10 | 19.04 | 49.03 | — | 24.04 | — | — | |
| BART-Long-Graphlength limit=10002021.10 | 18.99 | 49.24 | — | 23.97 | — | — | |
| BART-Long-Graph2023.08 | 18.99 | 49.24 | — | 23.97 | — | — | |
| LED+RELAX2023.08 | 18.86 | 47.23 | — | 25.03 | — | — | |
| PEGASUS2021.10 | 18.72 | 47.52 | — | 24.91 | — | — | |
| PEGASUS_LARGENumber of Parameters=470M2022.08 | 18.7 | — | — | — | — | — | |
| LongT5Model size=large, Input length=8k, Attention mechanism=TGlobal2021.12 | 18.44 | 47.18 | — | 24.18 | — | — | |
| LongT5_LARGENumber of Parameters=705M2022.08 | 18.4 | — | — | — | — | — | |
| GraphSum2023.08 | 17.56 | 45.87 | — | 23.39 | — | — | |
| TG-MultiSum2021.12 | 17.55 | 47.1 | — | 20.73 | — | — | |
| HiMAP2023.08 | 16.05 | 44.17 | — | 21.38 | — | — | |
| Hi-MAPlambda=0.52019.06 | 14.89 | 43.47 | 17.41 | — | — | — | |
| BARTBackbone=BART-large, Context=Literature Comparison2023.08 | 14.88 | 42.04 | — | 23.34 | — | — | |
| BARTBackbone=BART-large, Context=Local Baseline2023.08 | 14.69 | 42.16 | — | 23.51 | 60.7 | — | |
| PG-BRNN2019.06 | 14.19 | 42.8 | 16.75 | — | — | — | |
| CopyTransformer2019.06 | 14.03 | 43.57 | 17.37 | — | — | — | |
| PRIMERAzero-shot=true, output length limit=2562021.10 | 13.6 | 42 | — | 20.8 | — | — | |
| TextRank2019.06 | 13.1 | 38.44 | 13.5 | — | — | — | |
| PG-Original2019.06 | 12.91 | 41.85 | 16.46 | — | — | — | |
| LexRank2019.06 | 12.7 | 38.27 | 13.2 | — | — | — | |
| PG-MMR2019.06 | 12.36 | 40.55 | 15.87 | — | — | — | |
| BARTBackbone=BART-base, Context=Local Baseline2023.08 | 12.18 | 40.54 | — | 22.39 | 58.75 | — | |
| MMR2019.06 | 11.98 | 38.77 | 12.91 | — | — | — | |
| First-3strategy=First-k sentences/tokens2019.06 | 11.77 | 39.41 | 14.51 | — | — | — | |
| PEGASUSzero-shot=true, source=Zhang et al., 2020, output length limit=256, attention type=full-length attention (O(n^2)), pretrained on=single document datasets2021.10 | 10.5 | 36.5 | — | 18.7 | — | — | |
| First-2strategy=First-k sentences/tokens2019.06 | 10.17 | 35.99 | 12.06 | — | — | — | |
| PEGASUSzero-shot=true, source=our run, output length limit=2562021.10 | 10.1 | 32 | — | 16.7 | — | — | |
| First-1strategy=First-k sentences/tokens2019.06 | 7.25 | 26.83 | 6.46 | — | — | — | |
| BARTzero-shot=true, source=our run, output length limit=2562021.10 | 6.2 | 27.3 | — | 15.1 | — | — | |
| LEDzero-shot=true, source=our run, output length limit=2562021.10 | 3.7 | 17.3 | — | 10.4 | — | — | |
| Chain-of-Thought (CoT)Backbone=GPT-4.1, Shot-setting=2-shot2025.12 | — | — | — | — | 85.71 | 21.11 | |
| Cited Summarization (Cite)Backbone=GPT-4.1, Shot-setting=2-shot2025.12 | — | — | — | — | 83.86 | 18.51 | |
| Decomposition (Deco)Backbone=GPT-4.1, Shot-setting=2-shot2025.12 | — | — | — | — | 81.68 | 17.19 | |
| Extract-to-Abstract (E2A)Backbone=GPT-4.1, Shot-setting=2-shot2025.12 | — | — | — | — | 85.5 | 18.93 | |
| GPT-5Backbone=Large Reasoning Model, Shot-setting=2-shot2025.12 | — | — | — | — | 84.64 | 19.63 | |
| Iterative Refine (IR)Backbone=GPT-4.1, Shot-setting=2-shot2025.12 | — | — | — | — | 85.72 | 20.72 | |
| o1Backbone=Large Reasoning Model, Shot-setting=2-shot2025.12 | — | — | — | — | 83.53 | 17.67 | |
| o3Backbone=Large Reasoning Model, Shot-setting=2-shot2025.12 | — | — | — | — | 84.69 | 18.44 | |
| Plan-then-Write (Plan)Backbone=GPT-4.1, Shot-setting=2-shot2025.12 | — | — | — | — | 85.15 | 18.41 | |
| Question-Answer Guided (QAG)Backbone=GPT-4.1, Shot-setting=2-shot2025.12 | — | — | — | — | 84.63 | 20.52 | |
| Self-Consistency (SC)Backbone=GPT-4.1, Shot-setting=2-shot2025.12 | — | — | — | — | 84.77 | 19.51 | |
| VanillaBackbone=GPT-4.1, Shot-setting=2-shot2025.12 | — | — | — | — | 82.27 | 20.22 |