Faithfulness Evaluation on BookSum (test)
40.71SummaCGPT-5
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-5Model Category=LRM, Prompting=2-shot2025.12 | 40.71 | 78.62 | |
| o3Model Category=LRM, Prompting=2-shot2025.12 | 39.65 | 66.61 | |
| VanillaBase Model=GPT-4.1, Prompting=2-shot2025.12 | 39.04 | 66.56 | |
| Cited Summarization (Cite)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 38.81 | 67.53 | |
| Self-Consistency (SC)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 38.42 | 65.17 | |
| Chain-of-Thought (COT)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 38.31 | 67.03 | |
| o1Model Category=LRM, Prompting=2-shot2025.12 | 38.01 | 66.22 | |
| Extract-to-Abstract (E2A)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 37.77 | 65.07 | |
| Iterative Refine (IR)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 37.57 | 65.76 | |
| Decomposition (Deco)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 37.42 | 71.98 | |
| Plan-then-Write (Plan)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 36.98 | 70.21 | |
| Question-Answer Guided (QAG)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 34.95 | 69.07 |