Summarization on SAMSum (ROUGE-L, GPT-4o-Judge)
31.46ROUGE-LLongGuide
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| LongGuidebackbone=ChatGPT, shots=32025.06 | 31.46 | 7.72 | |
| LongGuidebackbone=Mistral-it (0.2), shots=32025.06 | 30.65 | 7.72 | |
| LongGuidebackbone=ChatGPT, shots=02025.06 | 30.47 | 7.59 | |
| LongGuidebackbone=Mistral-it (0.2), shots=02025.06 | 28.35 | 7.73 | |
| Mistral-it (0.2)shots=32025.06 | 27.13 | 7.66 | |
| APObackbone=Mistral-it (0.2), shots=32025.06 | 26.23 | 7.44 | |
| APObackbone=ChatGPT, shots=02025.06 | 25.05 | 7.45 | |
| APObackbone=ChatGPT, shots=32025.06 | 24.22 | 7.28 | |
| ChatGPTshots=02025.06 | 23.83 | 7.43 | |
| APObackbone=Mistral-it (0.2), shots=02025.06 | 23.77 | 7.31 | |
| ChatGPTshots=32025.06 | 22.21 | 7.32 | |
| Mistral-it (0.2)shots=02025.06 | 22.2 | 7.43 |