Abstractive Summarization on CNN/DailyMail full length F-1 (test)
41.69ROUGE-1DCA MLE+SEM+RL, mpgen, with-comm, with caa (m7)
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| DCA MLE+SEM+RL, mpgen, with-comm, with caa (m7)Loss=MLE+SEM+RL, Pointer=mpgen, Communication=with-comm, Attention=caa, Agents=32018.03 | 41.69 | 19.47 | 37.92 | — | |
| (m7) DCA MLE+SEM+RL, mpgen, with-comm, with caaModel Architecture=DCA, Loss Function=MLE+SEM+RL, Pointer Mechanism=mpgen, Communication Mechanism=with-comm, Attention Mechanism=caa, Number of Agents=3, Beam Search Width=52018.03 | 41.69 | 19.47 | 37.92 | — | |
| abstract-RLtraining=reinforcement learning2018.08 | 41.16 | 15.75 | 39.08 | — | |
| RL, with intra-attentionintra-attention=true, training=RL2018.03 | 41.16 | 15.75 | 39.08 | — | |
| RL, with intra-attentionAttention Mechanism=intra-decoder2018.03 | 41.16 | 15.75 | 39.08 | — | |
| DCA MLE+SEM, mpgen, with-comm, with caa (m6)Loss=MLE+SEM, Pointer=mpgen, Communication=with-comm, Attention=caa, Agents=32018.03 | 41.11 | 18.21 | 36.03 | — | |
| (m6) DCA MLE+SEM, mpgen, with-comm, with caaModel Architecture=DCA, Loss Function=MLE+SEM, Pointer Mechanism=mpgen, Communication Mechanism=with-comm, Attention Mechanism=caa, Number of Agents=3, Beam Search Width=52018.03 | 41.11 | 18.21 | 36.03 | — | |
| LATENTselection=top three sentences2018.08 | 41.05 | 18.77 | 37.54 | — | |
| EXTRACTarchitecture=LSTM, selection=top three sentences2018.08 | 40.62 | 18.45 | 37.14 | — | |
| LEAD3selection=first three leading sentences2018.08 | 40.34 | 17.7 | 36.57 | — | |
| EXTRACT-CNNarchitecture=CNN2018.08 | 40.11 | 17.52 | 36.39 | — | |
| REFRESHtraining=reinforcement learning2018.08 | 40 | 18.2 | 36.6 | — | |
| abstract-ML+RLtraining=reinforcement learning2018.08 | 39.87 | 15.82 | 36.9 | — | |
| ML+RL, with intra-attentionintra-attention=true, training=ML+RL2018.03 | 39.87 | 15.82 | 36.9 | — | |
| ML+RL, with intra-attentionAttention Mechanism=intra-decoder2018.03 | 39.87 | 15.82 | 36.9 | — | |
| Two-Layer Baseline (Pointer+Coverage) + Entailment GenerationMulti-task=Entailment Generation, Layers=22018.05 | 39.84 | 17.63 | 36.54 | 18.61 | |
| Two-Layer Baseline (Pointer+Coverage) + Entailment Gen. + Question Gen.Multi-task=Entailment and Question Generation, Layers=22018.05 | 39.81 | 17.64 | 36.54 | 18.54 | |
| controlled summarization with fixed values2018.03 | 39.75 | 17.29 | 36.54 | — | |
| controlled summarization with fixed values2018.03 | 39.75 | 17.29 | 36.54 | — | |
| Two-Layer Baseline (Pointer+Coverage) + Question GenerationMulti-task=Question Generation, Layers=22018.05 | 39.73 | 17.59 | 36.48 | 18.33 | |
| SummaRuNNerarchitecture=hierarchical recurrent neural networks2018.08 | 39.6 | 16.2 | 35.3 | — | |
| SummaRuNNer2018.03 | 39.6 | 16.2 | 35.3 | — | |
| SummaRuNNer2018.03 | 39.6 | 16.2 | 35.3 | — | |
| Two-Layer Baseline (Pointer+Coverage)Layers=2, Backbone=Pointer+Coverage2018.05 | 39.56 | 17.52 | 36.36 | 18.17 | |
| Pointer+CoverageSource=Reported in original paper (*)2018.05 | 39.53 | 17.28 | 36.38 | 18.72 | |
| pointer+coveragearchitecture=pointer generator variant2018.08 | 39.53 | 17.28 | 36.38 | — | |
| pointer generator + coveragecoverage=true2018.03 | 39.53 | 17.28 | 36.38 | — | |
| pointer generator + coverage2018.03 | 39.53 | 17.28 | 36.38 | — | |
| DCA MLE+SEM, mpgen, with-comm (m5)Loss=MLE+SEM, Pointer=mpgen, Communication=with-comm, Agents=32018.03 | 39.52 | 17.12 | 36.9 | — | |
| (m5) DCA MLE+SEM, mpgen, with-commModel Architecture=DCA, Loss Function=MLE+SEM, Pointer Mechanism=mpgen, Communication Mechanism=with-comm, Number of Agents=3, Beam Search Width=52018.03 | 39.52 | 17.12 | 36.9 | — | |
| LEAD3 (Nallapati et al., 2017)2018.08 | 39.2 | 15.7 | 35.5 | — | |
| Pointer+CoverageSource=Publicly released pretrained model (dagger)2018.05 | 38.82 | 16.81 | 35.71 | 18.14 | |
| graph-based attention2018.03 | 38.01 | 13.9 | 34 | — | |
| MLE+RL, pgen, no-comm (m3)Loss=MLE+RL, Pointer=pgen, Communication=no-comm, Agents=12018.03 | 38.01 | 16.43 | 35.49 | — | |
| graph-based attention2018.03 | 38.01 | 13.9 | 34 | — | |
| (m3) MLE+RL, pgen, no-comm (baseline-3)Loss Function=MLE+RL, Pointer Mechanism=pgen, Communication Mechanism=no-comm, Number of Agents=1, Beam Search Width=52018.03 | 38.01 | 16.43 | 35.49 | — | |
| DCA MLE+SEM, pgen, no-comm (m4)Loss=MLE+SEM, Pointer=pgen, Communication=no-comm, Agents=32018.03 | 37.45 | 15.9 | 34.56 | — | |
| (m4) DCA MLE+SEM, pgen, no-commModel Architecture=DCA, Loss Function=MLE+SEM, Pointer Mechanism=pgen, Communication Mechanism=no-comm, Number of Agents=3, Beam Search Width=52018.03 | 37.45 | 15.9 | 34.56 | — | |
| MLE+SEM, pgen, no-comm (m2)Loss=MLE+SEM, Pointer=pgen, Communication=no-comm, Agents=12018.03 | 36.9 | 15.02 | 33 | — | |
| (m2) MLE+SEM, pgen, no-comm (baseline-2)Loss Function=MLE+SEM, Pointer Mechanism=pgen, Communication Mechanism=no-comm, Number of Agents=1, Beam Search Width=52018.03 | 36.9 | 15.02 | 33 | — | |
| LATENT+COMPRESS2018.08 | 36.69 | 15.43 | 34.33 | — | |
| Pointer2018.05 | 36.44 | 15.66 | 33.42 | 15.35 | |
| pointer generator2018.03 | 36.44 | 15.66 | 33.42 | — | |
| pointer generator2018.03 | 36.44 | 15.66 | 33.42 | — | |
| MLE, pgen, no-comm (m1)Loss=MLE, Pointer=pgen, Communication=no-comm, Agents=12018.03 | 36.12 | 14.38 | 33.83 | — | |
| (m1) MLE, pgen, no-comm (baseline-1)Loss Function=MLE, Pointer Mechanism=pgen, Communication Mechanism=no-comm, Number of Agents=1, Beam Search Width=52018.03 | 36.12 | 14.38 | 33.83 | — | |
| abstractarchitecture=sequence-to-sequence2018.08 | 35.46 | 13.3 | 32.65 | — | |
| Seq2Seq (50k vocab)Vocabulary Size=50k2018.05 | 31.33 | 11.81 | 28.83 | 12.03 |