Binary Fact-checking on MediaSum
85.4Macro-F1Claude-3.7-Sonnet
Evaluation Results
| Method | Links | |
|---|---|---|
| Claude-3.7-SonnetModel Category=The State-of-the-Art LLMs, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | 85.4 | |
| o3Model Category=The State-of-the-Art LLMs, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | 82.9 | |
| InFi-Checker-QwenModel Category=Specialized Fact-Checking Models, Backbone=Qwen3-8B, Training=Fine-tuned2026.01 | 80.4 | |
| GPT-5Model Category=The State-of-the-Art LLMs, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | 80.2 | |
| FactCGModel Category=Specialized Fact-Checking Models2026.01 | 79.1 | |
| Qwen3-8BModel Category=The Open-Source Models, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | 77.7 | |
| GPT-4.1Model Category=The State-of-the-Art LLMs, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | 75.9 | |
| AlignScore-largeModel Category=Specialized Fact-Checking Models2026.01 | 75.8 | |
| MiniCheckModel Category=Specialized Fact-Checking Models2026.01 | 74.3 | |
| InFi-Checker-LlamaModel Category=Specialized Fact-Checking Models, Backbone=Llama-3.1-8B-Instruct, Training=Fine-tuned2026.01 | 73.5 | |
| GPT-4oModel Category=The State-of-the-Art LLMs, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | 71.5 | |
| ClearCheck (COT)Model Category=Specialized Fact-Checking Models2026.01 | 67.8 | |
| DeepSeek-V3.2-NoThinkModel Category=The State-of-the-Art LLMs, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | 65.5 | |
| Llama-3.1-8B-InstructModel Category=The Open-Source Models, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | 50.8 |