Faithfulness Evaluation on Reddit (test)
34.39SummaCGPT-5
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-5Model Category=LRM, Prompting=2-shot2025.12 | 34.39 | 21.41 | |
| VanillaBase Model=GPT-4.1, Prompting=2-shot2025.12 | 32.21 | 29.35 | |
| Chain-of-Thought (COT)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 27.5 | 49.92 | |
| o1Model Category=LRM, Prompting=2-shot2025.12 | 27.5 | 49.92 | |
| Question-Answer Guided (QAG)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 27.1 | 55.43 | |
| o3Model Category=LRM, Prompting=2-shot2025.12 | 26.54 | 35.54 | |
| Extract-to-Abstract (E2A)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 26.53 | 42.49 | |
| Self-Consistency (SC)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 25.33 | 50.8 | |
| Iterative Refine (IR)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 24.59 | 44.35 | |
| Plan-then-Write (Plan)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 24.31 | 55.33 | |
| Decomposition (Deco)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 24.26 | 56.46 | |
| Cited Summarization (Cite)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 24.17 | 50.1 |