Abstractive Summarization on ArXiv sampled (test)
23.65ROUGEIterative Refine (IR)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Iterative Refine (IR)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 23.65 | 84.85 | |
| Self-Consistency (SC)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 23.12 | 84.84 | |
| Cited Summarization (Cite)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 22.95 | 84.79 | |
| Extract-to-Abstract (E2A)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 22.68 | 84.52 | |
| Question-Answer Guided (QAG)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 22.23 | 84.2 | |
| VanillaShot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 20.97 | 83.76 | |
| Chain-of-Thought (CoT)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 20.39 | 83.59 | |
| o1Shot=0-shot, Model Category=Large Reasoning Models (LRM)2025.12 | 20.15 | 83.95 | |
| Decomposition (Deco)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 19.91 | 83.51 | |
| Plan-then-Write (Plan)Shot=0-shot, Model Category=LLM (GPT-4.1)2025.12 | 19.76 | 83.87 | |
| GPT-5Shot=0-shot, Model Category=Large Reasoning Models (LRM)2025.12 | 16.79 | 81.98 | |
| o3Shot=0-shot, Model Category=Large Reasoning Models (LRM)2025.12 | 16.61 | 81.71 |