Summarization on SciGen (test)
4.99Completeness ScoreChain-of-Thought (CoT)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Chain-of-Thought (CoT)Base Model=GPT-4.12025.12 | 4.99 | 3.93 | 4.98 | |
| Plan-then-Write (Plan)Base Model=GPT-4.12025.12 | 4.99 | 3.69 | 4.97 | |
| COTModel Category=LLM (GPT-4.1)2025.12 | 4.99 | 3.93 | 4.98 | |
| PlanModel Category=LLM (GPT-4.1)2025.12 | 4.99 | 3.69 | 4.97 | |
| o3Model Category=Large Reasoning Models (LRM)2025.12 | 4.98 | 4.4 | 4.98 | |
| o3Model Category=Large Reasoning Models (LRM)2025.12 | 4.98 | 4.4 | 4.98 | |
| VanillaBase Model=GPT-4.12025.12 | 4.97 | 4.61 | 4.97 | |
| GPT-5Model Category=Large Reasoning Models (LRM)2025.12 | 4.97 | 4.35 | 4.96 | |
| VanillaModel Category=LLM (GPT-4.1)2025.12 | 4.97 | 4.61 | 4.97 | |
| GPT-5Model Category=Large Reasoning Models (LRM)2025.12 | 4.97 | 4.35 | 4.96 | |
| Iterative Refine (IR)Base Model=GPT-4.12025.12 | 4.91 | 4.3 | 4.99 | |
| IRModel Category=LLM (GPT-4.1)2025.12 | 4.91 | 4.3 | 4.99 | |
| o1Model Category=Large Reasoning Models (LRM)2025.12 | 4.88 | 4.62 | 4.92 | |
| o1Model Category=Large Reasoning Models (LRM)2025.12 | 4.88 | 4.62 | 4.92 | |
| Decomposition (Deco)Base Model=GPT-4.12025.12 | 4.86 | 4.2 | 4.96 | |
| DecoModel Category=LLM (GPT-4.1)2025.12 | 4.86 | 4.2 | 4.96 | |
| Self-Consistency (SC)Base Model=GPT-4.12025.12 | 4.84 | 4.76 | 4.99 | |
| SCModel Category=LLM (GPT-4.1)2025.12 | 4.84 | 4.76 | 4.99 | |
| Cited Summarization (Cite)Base Model=GPT-4.12025.12 | 4.71 | 4.8 | 4.98 | |
| CiteModel Category=LLM (GPT-4.1)2025.12 | 4.71 | 4.8 | 4.98 | |
| Extract-to-Abstract (E2A)Base Model=GPT-4.12025.12 | 4.47 | 4.87 | 4.99 | |
| E2AModel Category=LLM (GPT-4.1)2025.12 | 4.47 | 4.87 | 4.99 | |
| Question-Answer Guided (QAG)Base Model=GPT-4.12025.12 | 4.38 | 3.92 | 4.99 | |
| QA-GModel Category=LLM (GPT-4.1)2025.12 | 4.38 | 3.92 | 4.99 |