Summarization Faithfulness on CNN/DM
52.85SummaCo3
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| o3System=o3, Model Class=Large Reasoning Models (LRM), Shots=0-shot2025.12 | 52.85 | 88.31 | — | |
| GPT-5System=GPT-5, Model Class=Large Reasoning Models (LRM), Shots=0-shot2025.12 | 48.59 | 91.57 | — | |
| VanillaSystem=Vanilla, Model Class=LLM (GPT-4.1), Shots=0-shot2025.12 | 45.49 | 83.49 | — | |
| Cited Summarization (Cite)System=Cite, Model Class=LLM (GPT-4.1), Shots=0-shot2025.12 | 43.91 | 11.92 | — | |
| Decomposition (Deco)System=Deco, Model Class=LLM (GPT-4.1), Shots=0-shot2025.12 | 43.67 | 66.06 | — | |
| Iterative Refine (IR)System=IR, Model Class=LLM (GPT-4.1), Shots=0-shot2025.12 | 40.99 | 8.53 | — | |
| Self-Consistency (SC)System=SC, Model Class=LLM (GPT-4.1), Shots=0-shot2025.12 | 40.44 | 6.39 | — | |
| Plan-then-Write (Plan)System=Plan, Model Class=LLM (GPT-4.1), Shots=0-shot2025.12 | 38.91 | 10.86 | — | |
| Extract-to-Abstract (E2A)System=E2A, Model Class=LLM (GPT-4.1), Shots=0-shot2025.12 | 33.58 | 43.24 | — | |
| Question-Answer Guided (QAG)System=QAG, Model Class=LLM (GPT-4.1), Shots=0-shot2025.12 | 31.59 | 40.7 | — | |
| o1System=o1, Model Class=Large Reasoning Models (LRM), Shots=0-shot2025.12 | 31.24 | 39.83 | — | |
| Chain-of-Thought (CoT)System=COT, Model Class=LLM (GPT-4.1), Shots=0-shot2025.12 | 30.94 | 42.67 | — | |
| FactCCYear=2020, Key Approach=BERT-based classification2026.06 | — | — | 74 | |
| SEAHORSEYear=2023, Key Approach=Question-based verification2026.06 | — | — | 85.3 |