Factuality Evaluation on AggreFact (FTSOTA)
70.5Balanced Accuracy (CNN-FTS)FENICE_GPT_claims
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| FENICE_GPT_claimsthresholding=single-threshold, claim_extractor=LLM-based (GPT)2024.03 | 70.5 | 72.8 | 71.6 | |
| QuestEvalthresholding=single-threshold2024.03 | 70.2 | 59.5 | 64.9 | |
| QAFactEvalthresholding=single-threshold2024.03 | 67.8 | 63.9 | 65.9 | |
| FENICE_T5_claimsthresholding=single-threshold, claim_extractor=knowledge-distilled (T5)2024.03 | 67.7 | 70.1 | 68.9 | |
| DAEthresholding=single-threshold2024.03 | 65.4 | 70.2 | 67.8 | |
| SummaC-ZSthresholding=single-threshold2024.03 | 64 | 56.4 | 60.2 | |
| MENLIthresholding=single-threshold2024.03 | 63.4 | 59 | 61.2 | |
| AlignScorethresholding=single-threshold2024.03 | 62.7 | 69.4 | 66.1 | |
| TrueTeacher-11Bthresholding=single-threshold, parameters=11B2024.03 | 62 | 74.9 | 68.4 | |
| SummaC-Convthresholding=single-threshold2024.03 | 61 | 65 | 63 | |
| ChatGPT-ZSthresholding=single-threshold, mode=zero-shot2024.03 | 56.3 | 62.7 | 59.5 | |
| ChatGPT-Starthresholding=single-threshold2024.03 | 56.3 | 57.8 | 57.1 | |
| ChatGPT-DAthresholding=single-threshold2024.03 | 53.7 | 54.9 | 54.3 | |
| ChatGPT-COTthresholding=single-threshold, mode=chain-of-thought2024.03 | 52.5 | 55.9 | 54.2 |