Veracity Assessment on FactCheck-Bench
91.3Macro-F1GPT-4.1
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GPT-4.1Model Category=The State-of-the-Art LLMs, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | 91.3 | — | — | |
| GPT-5Model Category=The State-of-the-Art LLMs, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | 90.4 | — | — | |
| FactCGModel Category=Specialized Fact-Checking Models2026.01 | 89 | — | — | |
| InFi-Checker-QwenModel Category=Specialized Fact-Checking Models, Backbone=Qwen3-8B, Training=Fine-tuned2026.01 | 88 | — | — | |
| DeepSeek-V3.2-NoThinkModel Category=The State-of-the-Art LLMs, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | 87.9 | — | — | |
| ClearCheck (COT)Model Category=Specialized Fact-Checking Models2026.01 | 87.9 | — | — | |
| Claude-3.7-SonnetModel Category=The State-of-the-Art LLMs, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | 86.9 | — | — | |
| o3Model Category=The State-of-the-Art LLMs, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | 86.9 | — | — | |
| MiniCheckModel Category=Specialized Fact-Checking Models2026.01 | 86.8 | — | — | |
| GPT-4oModel Category=The State-of-the-Art LLMs, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | 86 | — | — | |
| AlignScore-largeModel Category=Specialized Fact-Checking Models2026.01 | 83.7 | — | — | |
| InFi-Checker-LlamaModel Category=Specialized Fact-Checking Models, Backbone=Llama-3.1-8B-Instruct, Training=Fine-tuned2026.01 | 83.7 | — | — | |
| Qwen3-8BModel Category=The Open-Source Models, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | 78.4 | — | — | |
| MERMAIDLLM=GPT-4o2026.01 | 77 | 89 | 65 | |
| MERMAIDLLM=GPT-5-mini2026.01 | 77 | 88 | 66 | |
| FIRELLM=GPT-4o2026.01 | 76 | 85 | 66 | |
| MERMAIDLLM=Qwen-2.5-70B2026.01 | 75 | 89 | 61 | |
| FactCheck-GPTLLM=GPT-4o2026.01 | 74 | 83 | 65 | |
| SAFELLM=GPT-4o2026.01 | 74 | 84 | 65 | |
| MERMAIDLLM=OSS-120B2026.01 | 74 | 87 | 62 | |
| FacToolLLM=GPT-4o2026.01 | 73 | 82 | 64 | |
| MERMAIDLLM=Qwen-2.5-7B2026.01 | 72 | 84 | 60 | |
| Llama-3.1-8B-InstructModel Category=The Open-Source Models, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | 69.8 | — | — | |
| MERMAIDLLM=OSS-20B2026.01 | 67 | 81 | 53 | |
| MERMAIDLLM=LLaMA-3.1-70B-Inst2026.01 | 62 | 71 | 52 | |
| MERMAIDLLM=LLaMA-3.1-8B-Inst2026.01 | 54 | 62 | 45 |