Faithfulness Evaluation on CNN/DM (test)
35.56SummaCVanilla
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| VanillaBase Model=GPT-4.1, Prompting=2-shot2025.12 | 35.56 | 47.12 | |
| Extract-to-Abstract (E2A)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 32.85 | 37.28 | |
| o1Model Category=LRM, Prompting=2-shot2025.12 | 32.72 | 51.27 | |
| GPT-5Model Category=LRM, Prompting=2-shot2025.12 | 32.69 | 62.82 | |
| Question-Answer Guided (QAG)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 32 | 37.25 | |
| Cited Summarization (Cite)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 31.51 | 38.81 | |
| Chain-of-Thought (COT)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 30.94 | 42.67 | |
| Iterative Refine (IR)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 29.62 | 40.42 | |
| Self-Consistency (SC)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 29.43 | 40.67 | |
| o3Model Category=LRM, Prompting=2-shot2025.12 | 27.72 | 42.15 | |
| Decomposition (Deco)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 26.9 | 38.2 | |
| Plan-then-Write (Plan)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 26.72 | 35.59 |