Summarization Faithfulness on WikiHow
38.4SummaCVanilla
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| VanillaSystem=Vanilla, Model Class=LLM (GPT-4.1), Shots=0-shot2025.12 | 38.4 | 40.7 | |
| Cited Summarization (Cite)System=Cite, Model Class=LLM (GPT-4.1), Shots=0-shot2025.12 | 38.37 | 42.67 | |
| GPT-5System=GPT-5, Model Class=Large Reasoning Models (LRM), Shots=0-shot2025.12 | 37.14 | 42.34 | |
| Extract-to-Abstract (E2A)System=E2A, Model Class=LLM (GPT-4.1), Shots=0-shot2025.12 | 36.87 | 88.49 | |
| Decomposition (Deco)System=Deco, Model Class=LLM (GPT-4.1), Shots=0-shot2025.12 | 36.22 | 44.27 | |
| Self-Consistency (SC)System=SC, Model Class=LLM (GPT-4.1), Shots=0-shot2025.12 | 35.72 | 39.73 | |
| Iterative Refine (IR)System=IR, Model Class=LLM (GPT-4.1), Shots=0-shot2025.12 | 34.9 | 37.58 | |
| o1System=o1, Model Class=Large Reasoning Models (LRM), Shots=0-shot2025.12 | 34.67 | 94.46 | |
| o3System=o3, Model Class=Large Reasoning Models (LRM), Shots=0-shot2025.12 | 34.62 | 47.07 | |
| Chain-of-Thought (CoT)System=COT, Model Class=LLM (GPT-4.1), Shots=0-shot2025.12 | 34.06 | 94.17 | |
| Plan-then-Write (Plan)System=Plan, Model Class=LLM (GPT-4.1), Shots=0-shot2025.12 | 32.61 | 40.02 | |
| Question-Answer Guided (QAG)System=QAG, Model Class=LLM (GPT-4.1), Shots=0-shot2025.12 | 32.36 | 76.86 |