Faithfulness Evaluation on SAMSum (test)
29.58SummaCGPT-5
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-5Model Category=LRM, Prompting=2-shot2025.12 | 29.58 | 87.44 | |
| VanillaBase Model=GPT-4.1, Prompting=2-shot2025.12 | 29.03 | 85.05 | |
| Chain-of-Thought (COT)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 27.41 | 86.62 | |
| o1Model Category=LRM, Prompting=2-shot2025.12 | 27.41 | 86.62 | |
| Extract-to-Abstract (E2A)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 26.56 | 87.62 | |
| o3Model Category=LRM, Prompting=2-shot2025.12 | 25.59 | 85.87 | |
| Question-Answer Guided (QAG)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 25.34 | 86.88 | |
| Self-Consistency (SC)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 25.14 | 86.47 | |
| Iterative Refine (IR)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 25.11 | 88.71 | |
| Cited Summarization (Cite)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 24.68 | 86.83 | |
| Decomposition (Deco)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 24.36 | 88.02 | |
| Plan-then-Write (Plan)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 24.3 | 87.17 |