Faithfulness Evaluation on SciGen (test)
0.2822SummaC ScoreChain-of-Thought (COT)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Chain-of-Thought (COT)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 0.2822 | 0.9945 | |
| VanillaBase Model=GPT-4.1, Prompting=2-shot2025.12 | 0.2759 | 0.985 | |
| GPT-5Model Category=LRM, Prompting=2-shot2025.12 | 0.2653 | 0.9939 | |
| Extract-to-Abstract (E2A)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 0.2646 | 0.9871 | |
| o1Model Category=LRM, Prompting=2-shot2025.12 | 0.2646 | 0.9795 | |
| Question-Answer Guided (QAG)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 0.2562 | 0.9891 | |
| Decomposition (Deco)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 0.2504 | 0.9859 | |
| Cited Summarization (Cite)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 0.2466 | 0.9784 | |
| Self-Consistency (SC)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 0.2461 | 0.9844 | |
| Iterative Refine (IR)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 0.2447 | 0.987 | |
| o3Model Category=LRM, Prompting=2-shot2025.12 | 0.2384 | 0.9825 | |
| Plan-then-Write (Plan)Base Model=GPT-4.1, Prompting=2-shot2025.12 | 0.2346 | 0.9945 |