Abstractive Summarization on SciGen sampled (test)
18.42ROUGE ScoreIterative Refine (IR)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Iterative Refine (IR)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 18.42 | 85.34 | |
| Self-Consistency (SC)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 18.35 | 85.85 | |
| VanillaShot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 17.96 | 85.28 | |
| Decomposition (Deco)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 17.96 | 85.65 | |
| Question-Answer Guided (QAG)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 17.87 | 84.89 | |
| Extract-to-Abstract (E2A)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 17.84 | 85.78 | |
| Cited Summarization (Cite)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 16.93 | 85.79 | |
| Plan-then-Write (Plan)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 16.76 | 85.01 | |
| Chain-of-Thought (CoT)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 16.55 | 84.65 | |
| o1Shot=0-shot, Model Category=Large Reasoning Models (LRM)2025.12 | 15.21 | 81.16 | |
| o3Shot=0-shot, Model Category=Large Reasoning Models (LRM)2025.12 | 13.62 | 83.64 | |
| GPT-5Shot=0-shot, Model Category=Large Reasoning Models (LRM)2025.12 | 13.54 | 82.99 |