Fact-checking on ExpertQA
61.1Balanced AccuracyANCHOR
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ANCHORBackbone=Qwen2.5-72B, Direct Decision=false2026.05 | 61.1 | — | |
| GraphCheckBackbone=Llama3.3 70B2025.02 | 60.3 | — | |
| GPT-4Scale=>300B2025.02 | 59.6 | — | |
| AlignScore2025.02 | 59.3 | — | |
| CoTBackbone=Qwen2.5-72B, Direct Decision=false2026.05 | 59 | — | |
| OpenAI o1Scale=>300B2025.02 | 58.8 | — | |
| Claude 3.5-SonnetScale=>300B2025.02 | 58.8 | — | |
| DeepSeek-V3 671BScale=>300B2025.02 | 58.5 | — | |
| GPT-4oScale=>300B2025.02 | 58.3 | — | |
| BIRDBackbone=Qwen2.5-72B, Direct Decision=false2026.05 | 58.2 | — | |
| ACUEval2025.02 | 57.5 | — | |
| MiniCheck2025.02 | 57.4 | — | |
| GraphCheckBackbone=Qwen 72B2025.02 | 57.2 | — | |
| VanillaBackbone=DeepSeek-V3-671B, Direct Decision=true2026.05 | 56.2 | — | |
| GraphEval2025.02 | 56 | — | |
| CoTBackbone=Qwen2.5-72B, Direct Decision=true2026.05 | 55.2 | — | |
| Llama3.3 70B2025.02 | 54.3 | — | |
| Qwen2.5 72B2025.02 | 54.1 | — | |
| CoTBackbone=DeepSeek-V3-671B, Direct Decision=false2026.05 | 54.1 | — | |
| VanillaBackbone=DeepSeek-V3-671B, Direct Decision=false2026.05 | 53.8 | — | |
| Qwen2.5 7B2025.02 | 53.6 | — | |
| VanillaBackbone=Qwen2.5-72B, Direct Decision=true2026.05 | 53 | — | |
| VanillaBackbone=Qwen2.5-72B, Direct Decision=false2026.05 | 51.7 | — | |
| Llama3 8B2025.02 | 51.3 | — | |
| CoTBackbone=DeepSeek-V3-671B, Direct Decision=true2026.05 | 49.8 | — | |
| AlignScore-largeModel Category=Specialized Fact-Checking Models2026.01 | — | 75 | |
| Claude-3.7-SonnetModel Category=The State-of-the-Art LLMs, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | — | 74.4 | |
| ClearCheck (COT)Model Category=Specialized Fact-Checking Models2026.01 | — | 72.7 | |
| DeepSeek-V3.2-NoThinkModel Category=The State-of-the-Art LLMs, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | — | 74.4 | |
| FactCGModel Category=Specialized Fact-Checking Models2026.01 | — | 75.3 | |
| GPT-4.1Model Category=The State-of-the-Art LLMs, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | — | 80.3 | |
| GPT-4oModel Category=The State-of-the-Art LLMs, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | — | 68.3 | |
| GPT-5Model Category=The State-of-the-Art LLMs, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | — | 75.9 | |
| InFi-Checker-LlamaModel Category=Specialized Fact-Checking Models, Backbone=Llama-3.1-8B-Instruct, Training=Fine-tuned2026.01 | — | 78.3 | |
| InFi-Checker-QwenModel Category=Specialized Fact-Checking Models, Backbone=Qwen3-8B, Training=Fine-tuned2026.01 | — | 75.7 | |
| Llama-3.1-8B-InstructModel Category=The Open-Source Models, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | — | 49.8 | |
| MiniCheckModel Category=Specialized Fact-Checking Models2026.01 | — | 72.9 | |
| o3Model Category=The State-of-the-Art LLMs, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | — | 79.6 | |
| Qwen3-8BModel Category=The Open-Source Models, Prompting Format=InFi-Check reasoning format, Number of Shots=zero/one-shot2026.01 | — | 55.6 |