Abstractive Summarization on XSUM (test)
40.4ROUGE-LBRIO-Mul
Evaluation Results
| Method | Links | ||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| BRIO-MulEvaluation Source=Original Paper2023.05 | 40.4 | — | — | — | — | — | — | — | — | — | 49.07 | 25.59 | — | — | — | — | |
| BRIO-MulEvaluation Source=Our evaluation script2023.05 | 40.16 | — | — | — | — | — | — | — | — | — | 48.74 | 25.38 | 92.6 | — | — | — | |
| SummaRerankerEvaluation Source=Original Paper2023.05 | 40 | — | — | — | — | — | — | — | — | — | 48.12 | 24.95 | 92.14 | — | — | — | |
| BRIO-CtrEvaluation Source=Our evaluation script2023.05 | 39.96 | — | — | — | — | — | — | — | — | — | 48.12 | 25.24 | 91.72 | — | — | — | |
| BRIO-CtrEvaluation Source=Original Paper2023.05 | 39.84 | — | — | — | — | — | — | — | — | — | 48.13 | 25.13 | — | — | — | — | |
| SimCLSEvaluation Source=Original Paper2023.05 | 39.44 | — | — | — | — | — | — | — | — | — | 47.61 | 24.57 | — | — | — | — | |
| SimCLSEvaluation Source=Our evaluation script2023.05 | 39.31 | — | — | — | — | — | — | — | — | — | 47.37 | 24.49 | 91.48 | — | — | — | |
| PegasusEvaluation Source=Original Paper2023.05 | 39.25 | — | — | — | — | — | — | — | — | — | 47.21 | 24.56 | — | — | — | — | |
| BalSumEvaluation Source=Our evaluation script2023.05 | 39.09 | — | — | — | — | — | — | — | — | — | 47.17 | 24.23 | 91.48 | — | — | — | |
| PegasusEvaluation Source=Our evaluation script2023.05 | 39.07 | — | — | — | — | — | — | — | — | — | 46.82 | 24.44 | 91.93 | — | — | — | |
| BARTEvaluation Source=Original Paper2023.05 | 37.25 | — | — | — | — | — | — | — | — | — | 45.14 | 22.27 | — | — | — | — | |
| Z-Code++ Largetraining_examples=1000, pretraining=Two-phase2022.08 | 33.6 | — | — | — | — | — | — | — | — | — | 41.9 | 18.9 | — | — | — | — | |
| PEGASUS-Largetraining_examples=10002022.08 | 33.3 | — | — | — | — | — | — | — | — | — | 41.6 | 18.2 | — | — | — | — | |
| Z-Code++ Large (Phase 1)training_examples=1000, pretraining=Phase 1 only2022.08 | 32.8 | — | — | — | — | — | — | — | — | — | 40.9 | 17.3 | — | — | — | — | |
| PEGASUS-Largetraining_examples=1002022.08 | 31.3 | — | — | — | — | — | — | — | — | — | 39.07 | 16.4 | — | — | — | — | |
| Z-Code++ Largetraining_examples=100, pretraining=Two-phase2022.08 | 30 | — | — | — | — | — | — | — | — | — | 40.6 | 17.5 | — | — | — | — | |
| Z-Code++ Largetraining_examples=10, pretraining=Two-phase2022.08 | 29.1 | — | — | — | — | — | — | — | — | — | 37.4 | 14 | — | — | — | — | |
| Z-Code++ Largetraining_examples=0, zero-shot=true, pretraining=Two-phase2022.08 | 28.6 | — | — | — | — | — | — | — | — | — | 36.6 | 13.7 | — | — | — | — | |
| Z-Code++ Large (Phase 1)training_examples=100, pretraining=Phase 1 only2022.08 | 27.5 | — | — | — | — | — | — | — | — | — | 35.3 | 12.3 | — | — | — | — | |
| GPSTmedium#param.=2.12024.03 | 25.58 | — | — | — | — | — | — | — | — | — | 31.96 | 11.31 | — | 22.95 | — | — | |
| GPT-2medium_2.5#param.=2.12024.03 | 25.35 | — | — | — | — | — | — | — | — | — | 31.95 | 11.17 | — | 22.82 | — | — | |
| GPT-2medium#param.=2.02024.03 | 25.28 | — | — | — | — | — | — | — | — | — | 31.91 | 11.11 | — | 22.76 | — | — | |
| GPSTmedium w/o sync#param.=2.12024.03 | 25.16 | — | — | — | — | — | — | — | — | — | 31.66 | 10.91 | — | 22.58 | — | — | |
| T5-Largetraining_examples=10002022.08 | 23.8 | — | — | — | — | — | — | — | — | — | 31.2 | 9.4 | — | — | — | — | |
| GPSTsmall#param.=1.052024.03 | 23.7 | — | — | — | — | — | — | — | — | — | 29.86 | 9.51 | — | 21.02 | — | — | |
| GPT-2small_1.3#param.=1.12024.03 | 23.62 | — | — | — | — | — | — | — | — | — | 29.84 | 9.46 | — | 20.97 | — | — | |
| GPT-2small#param.=1.02024.03 | 23.56 | — | — | — | — | — | — | — | — | — | 29.78 | 9.43 | — | 20.92 | — | — | |
| GPSTsmall-w/o sync#param.=1.052024.03 | 23.2 | — | — | — | — | — | — | — | — | — | 29.44 | 9.09 | — | 20.58 | — | — | |
| CoRectBackbone=LLaMA-3-8b2026.02 | 20.04 | — | — | — | — | — | — | — | — | — | — | — | 87.3 | — | — | 83.3 | |
| T5-Largetraining_examples=1002022.08 | 17 | — | — | — | — | — | — | — | — | — | 21.5 | 5.5 | — | — | — | — | |
| GreedyBackbone=LLaMA-3-8b2026.02 | 16.42 | — | — | — | — | — | — | — | — | — | — | — | 86.56 | — | — | 77.65 | |
| AdaCADBackbone=LLaMA-3-8b2026.02 | 15.81 | — | — | — | — | — | — | — | — | — | — | — | 86.43 | — | — | 82.02 | |
| COIECDBackbone=LLaMA-3-8b2026.02 | 15.77 | — | — | — | — | — | — | — | — | — | — | — | 86.48 | — | — | 81.06 | |
| CADBackbone=LLaMA-3-8b2026.02 | 15.25 | — | — | — | — | — | — | — | — | — | — | — | 84.3 | — | — | 67 | |
| Self-training + SummScoreDecoding=Beam search, Beams=20, Re-ranking=SummScore2022.12 | 14.93 | — | — | — | — | — | — | — | — | — | 20.02 | 2.84 | 86.23 | — | — | — | |
| PEGASUS (ours) + SummScoreDecoding=Beam search, Beams=20, Re-ranking=SummScore2022.12 | 14.71 | — | — | — | — | — | — | — | — | — | 19.62 | 3.02 | 85.92 | — | — | — | |
| Self-trainingDecoding=Beam search, Beams=202022.12 | 14.18 | — | — | — | — | — | — | — | — | — | 19.33 | 2.76 | 86.03 | — | — | — | |
| PEGASUS-Largetraining_examples=102022.08 | 14.02 | — | — | — | — | — | — | — | — | — | 19.4 | 3.5 | — | — | — | — | |
| PEGASUS (ours)Decoding=Beam search, Beams=202022.12 | 13.85 | — | — | — | — | — | — | — | — | — | 18.77 | 2.86 | 85.66 | — | — | — | |
| PEGASUS-Largetraining_examples=0, zero-shot=true2022.08 | 12.7 | — | — | — | — | — | — | — | — | — | 19.3 | 3 | — | — | — | — | |
| Z-Code++ Large (Phase 1)training_examples=10, pretraining=Phase 1 only2022.08 | 12.6 | — | — | — | — | — | — | — | — | — | 16.7 | 2.1 | — | — | — | — | |
| T5-Largetraining_examples=102022.08 | 10 | — | — | — | — | — | — | — | — | — | 13.2 | 2.5 | — | — | — | — | |
| T5-Largetraining_examples=0, zero-shot=true2022.08 | 9.8 | — | — | — | — | — | — | — | — | — | 12.8 | 2.3 | — | — | — | — | |
| Z-Code++ Large (Phase 1)training_examples=0, zero-shot=true, pretraining=Phase 1 only2022.08 | 3.7 | — | — | — | — | — | — | — | — | — | 3.6 | 0.1 | — | — | — | — | |
| PEGASUSSetting=Unsupervised2022.12 | — | 96.78 | — | — | — | — | — | — | — | — | — | — | — | — | 4.54 | — | |
| QUALS-CONSEQbaseline=BART-large MLE, sample_size=100, evaluation_method=human preference (majority vote of 3)2021.05 | — | 18 | 9 | 73 | 22 | 9 | 69 | 4 | 2 | 94 | — | — | — | — | — | — | |
| SummScoreSetting=Unsupervised2022.12 | — | 97.53 | — | — | — | — | — | — | — | — | — | — | — | — | 4.64 | — |