Faithfulness Evaluation on ArXiv (test)
53.58SummaCo3
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| o3Model Category=LRM, Prompting=2-shot2025.12 | 53.58 | 85.22 | |
| GPT-5Model Category=LRM, Prompting=2-shot2025.12 | 45.9 | 85.55 | |
| o1Model Category=LRM, Prompting=2-shot2025.12 | 44.24 | 27.61 | |
| VanillaBase Model=GPT-4.1, Prompting=2-shot2025.12 | 43.36 | 26.07 | |
| Extract-to-Abstract (E2A)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 42.96 | 8.68 | |
| Cited Summarization (Cite)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 41.79 | 5.99 | |
| Self-Consistency (SC)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 40.68 | 8 | |
| Chain-of-Thought (COT)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 40.54 | 7.46 | |
| Decomposition (Deco)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 40.37 | 17.88 | |
| Question-Answer Guided (QAG)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 40.37 | 26.94 | |
| Iterative Refine (IR)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 39.87 | 5.65 | |
| Plan-then-Write (Plan)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 39.09 | 9.52 |