Faithfulness Evaluation on Multi-News (test)
38.5SummaCGPT-5
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-5Model Category=LRM, Prompting=2-shot2025.12 | 38.5 | 50.65 | |
| Extract-to-Abstract (E2A)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 37.31 | 43.38 | |
| o1Model Category=LRM, Prompting=2-shot2025.12 | 37.08 | 49.02 | |
| Chain-of-Thought (COT)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 37.07 | 39.78 | |
| VanillaBase Model=GPT-4.1, Prompting=2-shot2025.12 | 36.96 | 18.87 | |
| Cited Summarization (Cite)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 36.05 | 38.2 | |
| Question-Answer Guided (QAG)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 35.86 | 40.98 | |
| Self-Consistency (SC)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 35.39 | 40.53 | |
| Iterative Refine (IR)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 34.79 | 40.57 | |
| o3Model Category=LRM, Prompting=2-shot2025.12 | 34.12 | 37.48 | |
| Decomposition (Deco)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 33.68 | 37.27 | |
| Plan-then-Write (Plan)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 33.01 | 41.17 |