Fact-checking on FEVEROUS (test)
74.72Macro F1Trification
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| TrificationModel Category=Our, Method Components=Full2025.11 | 74.72 | — | — | |
| PACARModel Category=III2025.11 | 72.61 | — | — | |
| TrificationModel Category=Our, Method Components=w/o REFINE2025.11 | 71.3 | — | — | |
| TrificationModel Category=Our, Method Components=w/o REFINE and THINK2025.11 | 70.58 | — | — | |
| LoCaLModel Category=IV2025.11 | 68.22 | — | — | |
| ProgramFCModel Category=III, Number of programs (N)=52025.11 | 68.06 | — | — | |
| ProgramFCModel Category=III, Number of programs (N)=12025.11 | 67.8 | — | — | |
| SearChainModel Category=IV2025.11 | 66.69 | — | — | |
| Flan-T5Model Category=II2025.11 | 63.73 | — | — | |
| CodexModel Category=II2025.11 | 62.58 | — | — | |
| DeBERTaV3-NLIModel Category=I2025.11 | 58.81 | — | — | |
| RoBERTa-NLIModel Category=I2025.11 | 57.8 | — | — | |
| MUTIVERSModel Category=I2025.11 | 56.61 | — | — | |
| causal explanation-based verdict prediction systemKnowledge Source=LLMs, Evaluation Protocol=Tolerant2025.12 | 56 | 52 | 62 | |
| causal explanation-based verdict prediction systemKnowledge Source=Common Sense, Evaluation Protocol=Tolerant2025.12 | 56 | 52 | 62 | |
| ChatGPTModel Category=II2025.11 | 55.72 | — | — | |
| LisT5Model Category=I2025.11 | 54.15 | — | — | |
| BERT-FCModel Category=I2025.11 | 51.67 | — | — | |
| causal explanation-based verdict prediction systemKnowledge Source=LLMs, Evaluation Protocol=Strict2025.12 | 47 | 50 | 44 | |
| causal explanation-based verdict prediction systemKnowledge Source=Common Sense, Evaluation Protocol=Strict2025.12 | 47 | 51 | 44 |