Binary Fact-checking on REVEAL
93.7Macro-F1GPT-5
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-5Model Category=The State-of-the-Art LLMs, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | 93.7 | |
| GPT-4.1Model Category=The State-of-the-Art LLMs, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | 93.2 | |
| o3Model Category=The State-of-the-Art LLMs, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | 92.2 | |
| AlignScore-largeModel Category=Specialized Fact-Checking Models2026.01 | 92.2 | |
| DeepSeek-V3.2-NoThinkModel Category=The State-of-the-Art LLMs, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | 91 | |
| MiniCheckModel Category=Specialized Fact-Checking Models2026.01 | 91 | |
| FactCGModel Category=Specialized Fact-Checking Models2026.01 | 90 | |
| InFi-Checker-QwenModel Category=Specialized Fact-Checking Models, Backbone=Qwen3-8B, Training=Fine-tuned2026.01 | 90 | |
| Claude-3.7-SonnetModel Category=The State-of-the-Art LLMs, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | 88 | |
| InFi-Checker-LlamaModel Category=Specialized Fact-Checking Models, Backbone=Llama-3.1-8B-Instruct, Training=Fine-tuned2026.01 | 87.7 | |
| ClearCheck (COT)Model Category=Specialized Fact-Checking Models2026.01 | 87 | |
| GPT-4oModel Category=The State-of-the-Art LLMs, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | 86.9 | |
| Qwen3-8BModel Category=The Open-Source Models, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | 83.2 | |
| Llama-3.1-8B-InstructModel Category=The Open-Source Models, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | 78.2 |