Binary Fact-checking on Claim Verify
0.896Macro F1InFi-Checker-Qwen
Evaluation Results
| Method | Links | |
|---|---|---|
| InFi-Checker-QwenModel Category=Specialized Fact-Checking Models, Backbone=Qwen3-8B, Training=Fine-tuned2026.01 | 0.896 | |
| GPT-5Model Category=The State-of-the-Art LLMs, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | 0.877 | |
| MiniCheckModel Category=Specialized Fact-Checking Models2026.01 | 0.856 | |
| ClearCheck (COT)Model Category=Specialized Fact-Checking Models2026.01 | 0.854 | |
| Claude-3.7-SonnetModel Category=The State-of-the-Art LLMs, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | 0.837 | |
| o3Model Category=The State-of-the-Art LLMs, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | 0.833 | |
| GPT-4.1Model Category=The State-of-the-Art LLMs, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | 0.816 | |
| AlignScore-largeModel Category=Specialized Fact-Checking Models2026.01 | 0.798 | |
| GPT-4oModel Category=The State-of-the-Art LLMs, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | 0.783 | |
| FactCGModel Category=Specialized Fact-Checking Models2026.01 | 0.762 | |
| InFi-Checker-LlamaModel Category=Specialized Fact-Checking Models, Backbone=Llama-3.1-8B-Instruct, Training=Fine-tuned2026.01 | 0.759 | |
| DeepSeek-V3.2-NoThinkModel Category=The State-of-the-Art LLMs, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | 0.754 | |
| Qwen3-8BModel Category=The Open-Source Models, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | 0.661 | |
| Llama-3.1-8B-InstructModel Category=The Open-Source Models, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | 0.636 |