Summarization on MovieSum (test)
33.92ROUGE-1TextRank
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| TextRankmerging_strategy=without merging2025.01 | 33.92 | 4.62 | 16.25 | 46.82 | 49.48 | 48.1 | 60.23 | — | — | |
| Agent-as-a-Judgemerging_strategy=hierarchical merging, iteration=3rd, backbone=GPT-4o-mini2025.01 | 31.31 | 8.81 | 18.62 | 59.36 | 59.31 | 59.32 | 94.59 | — | — | |
| Agent-as-a-Judgemerging_strategy=hierarchical merging, iteration=2nd, backbone=GPT-4o-mini2025.01 | 30.98 | 8.75 | 18.61 | 59.33 | 59.3 | 59.3 | 92.04 | — | — | |
| Agent-as-a-Judgemerging_strategy=hierarchical merging, iteration=1st, backbone=GPT-4o-mini2025.01 | 30.36 | 8.74 | 18.55 | 59.26 | 59.3 | 59.27 | 86.92 | — | — | |
| GPT-4o-minimerging_strategy=hierarchical merging2025.01 | 29.26 | 8.72 | 17.88 | 59.11 | 59.29 | 59.19 | 80.56 | — | — | |
| LongT5merging_strategy=without merging2025.01 | 20.18 | 1.99 | 13.83 | 44.58 | 44.28 | 44.36 | 74.01 | — | — | |
| LEDmerging_strategy=without merging2025.01 | 2.8 | 0.28 | 0.28 | 32.64 | 23.82 | 27.32 | 22.24 | — | — | |
| Description Onlybackbone=LED-Large, size=459M, category=Extractive-to-Abstractive2025.05 | — | — | — | — | — | — | — | 58.92 | — | |
| GPT4omode=Zero-Shot, category=Long Context Modeling2025.05 | — | — | — | — | — | — | — | 52.8 | — | |
| HM-SRbackbone=GPT4o-mini, category=Multi-LLM Agent2025.05 | — | — | — | — | — | — | — | 59.32 | — | |
| Mistral Largemode=Zero-Shot, size=123B, category=Long Context Modeling2025.05 | — | — | — | — | — | — | — | 55.5 | — | |
| NEXUSSUMbackbone=Mistral Large, size=123B, category=Multi-LLM Agent2025.05 | — | — | — | — | — | — | — | 63.53 | — | |
| Two-Stage Heuristicbackbone=LED-Large, size=459M, category=Extractive-to-Abstractive2025.05 | — | — | — | — | — | — | — | 58.54 | — |