Fact-checking Explanation Generation on Combined Datasets (Overall)
73Helpfulness ScoreCLUE
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| CLUEBackbone=Qwen2.5-14B-Instruct2025.05 | 73 | 72.1 | 76.2 | 72.2 | 67.8 | |
| PromptBaselineBackbone=Qwen2.5-14B-Instruct2025.05 | 28.1 | 29 | 25.2 | 27.5 | 32.5 |