Fact Checking on RealHitBench
70.91Exact MatchDeepSeek-R1
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| DeepSeek-R1Input=Text2025.06 | 70.91 | 79.45 | |
| DeepSeek-R1Prompting Strategy=Direct2026.04 | 70.91 | 79.45 | |
| Qwen3-Coder-480B w/ SpreadsheetAgentTool=Python2026.04 | 69.1 | 77.18 | |
| TreeThinkerModel Backbone=Llama3.3-70B-Instruct, Prompting Strategy=Explicit Reasoning2026.04 | 67.3 | 73.06 | |
| Llama3.3-70B-InstructPrompting Strategy=Explicit Reasoning2026.04 | 67.13 | 73.14 | |
| TreeThinkerModel Backbone=Qwen2.5-72B-Instruct, Prompting Strategy=Explicit Reasoning2026.04 | 66.8 | 74.03 | |
| SpreadsheetAgentModel Backbone=Qwen2.5-72B-Instruct, Prompting Strategy=Explicit Reasoning2026.04 | 66.64 | 73.44 | |
| Qwen3-Coder-480BTool=Python2026.04 | 66.47 | 74.46 | |
| QwQ-32BModel Scale Group=Larger2025.12 | 66.31 | — | |
| GPT4o(TreeThinker)Input=Image+Text2025.06 | 65.82 | 73.32 | |
| SpreadsheetAgentModel Backbone=Llama3.3-70B-Instruct, Prompting Strategy=Explicit Reasoning2026.04 | 65.74 | 71.94 | |
| Qwen2.5-72B-InstructPrompting Strategy=Explicit Reasoning2026.04 | 65.74 | 72.93 | |
| DeepSeek-V3Model Scale Group=Larger2025.12 | 65.08 | — | |
| Qwen3-32BModel Scale Group=Larger2025.12 | 65 | — | |
| GPT4o(TreeThinker)Input=Text2025.06 | 64.5 | 72.41 | |
| GPT-OSS-120B w/ SpreadsheetAgentTool=Python2026.04 | 64.09 | 69.22 | |
| TableGPT-R1-8BModel Scale Group=Comparable2025.12 | 63.85 | — | |
| Qwen3-235B w/ SpreadsheetAgentTool=Python2026.04 | 62.7 | 68.55 | |
| GPT4oInput=Image+Text2025.06 | 62.45 | 69.01 | |
| Qwen3-14BModel Scale Group=Larger2025.12 | 62.36 | — | |
| Gemini1.5-proInput=Image+Text2025.06 | 62.04 | 68.08 | |
| GPT4oInput=Text2025.06 | 60.31 | 68.97 | |
| GPT4oPrompting Strategy=Direct2026.04 | 60.31 | 68.97 | |
| Qwen3-72BModel Scale Group=Larger2025.12 | 60.23 | — | |
| Doubao-1.5-pro-32kInput Modality=Text2025.06 | 59.65 | 66.34 | |
| Gemini1.5-proInput=Text2025.06 | 59.08 | 66.14 | |
| Gemini1.5-proPrompting Strategy=Direct2026.04 | 59.08 | 66.14 | |
| Qwen3-8BModel Scale Group=Comparable2025.12 | 58.83 | — | |
| DTRBackbone=DeepSeek-v32026.03 | 58.22 | 64.47 | |
| GPT-OSS-120BTool=Python2026.04 | 58.18 | 62.75 | |
| Qwen3-Coder-480B w/ TreeThinkerTool=Python2026.04 | 57.85 | 65.87 | |
| Qwen3-235B w/ TreeThinkerTool=Python2026.04 | 57.27 | 62.54 | |
| DeepSeek-v32026.03 | 57.21 | 53.42 | |
| Qwen3-235BTool=Python2026.04 | 56.94 | 62.6 | |
| Qwen-PlusModel Scale Group=Larger2025.12 | 56.53 | — | |
| GPT-OSS-120B w/ TreeThinkerTool=Python2026.04 | 56.53 | 61.14 | |
| Llama3.2-90B-Vision-InstructInput=Image+Text2025.06 | 55.6 | 65.19 | |
| GPT-4oModel Scale Group=Larger2025.12 | 55.22 | — | |
| Llama3.3-70B-InstructInput=Text2025.06 | 53.08 | 64.53 | |
| Llama3.3-70B-InstructPrompting Strategy=Direct2026.04 | 53.08 | 64.53 | |
| Llama3.2-90B-Vision-InstructInput=Text2025.06 | 52.1 | 61.84 | |
| Qwen2.5-72B-InstructInput=Text2025.06 | 51.93 | 62.15 | |
| Qwen2.5-72B-InstructPrompting Strategy=Direct2026.04 | 51.93 | 62.15 | |
| SpreadsheetAgentModel Backbone=Qwen2.5-7B-Instruct, Prompting Strategy=Explicit Reasoning2026.04 | 51.27 | 57.56 | |
| TreeThinkerModel Backbone=Qwen2.5-7B-Instruct, Prompting Strategy=Explicit Reasoning2026.04 | 51.19 | 57.45 | |
| Gemini1.5-proInput=Image2025.06 | 50.37 | 57.62 | |
| GPT-OSS-20B w/ SpreadsheetAgentTool=Python2026.04 | 48.73 | 53.64 | |
| Qwen2.5-7B-InstructPrompting Strategy=Explicit Reasoning2026.04 | 48.56 | 54.7 | |
| Code LoopBackbone=DeepSeek-v32026.03 | 48.19 | 56.49 | |
| Qwen3-30B w/ SpreadsheetAgentTool=Python2026.04 | 46.18 | 52.44 | |
| TableGPT2-7BInput=Text2025.06 | 46.1 | 53.8 | |
| TableGPT2-7B2026.03 | 46.1 | 53.8 | |
| GPT-OSS-20B w/ TreeThinkerTool=Python2026.04 | 45.11 | 50.12 | |
| GPT4o(TreeThinker)Input=Image2025.06 | 44.13 | 52.41 | |
| GPT4oInput=Image2025.06 | 43.39 | 51.87 | |
| GPT4o2026.03 | 43.39 | 51.87 | |
| TableGPT2-7BModel Scale Group=Comparable2025.12 | 43.06 | — | |
| GPT-OSS-20BTool=Python2026.04 | 42.81 | 48.54 | |
| TableLLM-7B2026.03 | 38.25 | 44.12 | |
| Qwen3-30B w/ TreeThinkerTool=Python2026.04 | 37.63 | 45.03 | |
| Qwen3-30BTool=Python2026.04 | 36.24 | 42.31 | |
| SpreadsheetAgentModel Backbone=Llama3.1-8B-Instruct, Prompting Strategy=Explicit Reasoning2026.04 | 34.51 | 47.68 | |
| TableLLM-Qwen2-7BInput=Text2025.06 | 33.53 | 40.41 | |
| TableLLMModel Scale Group=Comparable2025.12 | 33.44 | — | |
| TableLLM-Llama3.1-8BInput=Text2025.06 | 33.44 | 39.49 | |
| Llama3.1-8B-InstructPrompting Strategy=Explicit Reasoning2026.04 | 33.2 | 44.15 | |
| Qwen2-VL-7B-InstructInput Modality=Image+Text2025.06 | 32.79 | 45.19 | |
| TreeThinkerModel Backbone=Llama3.1-8B-Instruct, Prompting Strategy=Explicit Reasoning2026.04 | 32.37 | 44.93 | |
| DTRBackbone=Qwen3-4B2026.03 | 30.42 | 37.44 | |
| Llama-3.1-8BModel Scale Group=Comparable2025.12 | 30.32 | — | |
| Llama3.1-8B-InstructInput=Text2025.06 | 30.32 | 44.93 | |
| Llama3.1-8B-InstructPrompting Strategy=Direct2026.04 | 30.32 | 44.93 | |
| QwQ-32BInput Modality=Text2025.06 | 28.95 | 43.13 | |
| Qwen2-VL-7B-InstructInput Modality=Image2025.06 | 28.5 | 38.34 | |
| DeepSeek-R1-Distill-Qwen-7BInput Modality=Text2025.06 | 28.1 | 34.53 | |
| StructGPT2026.03 | 25.4 | 32.15 | |
| Llama3.2-90B-Vision-InstructInput=Image2025.06 | 23.75 | 36.08 | |
| DeepSeek-R1-Distill-Llama-8BInput Modality=Text2025.06 | 22.43 | 33.29 | |
| Llama3.2-11B-Vision-InstructInput=Image+Text2025.06 | 19.39 | 32.54 | |
| Qwen2.5-7B-InstructInput=Text2025.06 | 18.65 | 38.39 | |
| Qwen2.5-7B-InstructPrompting Strategy=Direct2026.04 | 18.65 | 38.39 | |
| Llama3.2-11B-Vision-InstructInput=Text2025.06 | 18.41 | 38.44 | |
| DTRBackbone=Qwen3-1.7B2026.03 | 17.93 | 23.5 | |
| Code LoopBackbone=Qwen3-4B2026.03 | 16.61 | 19.76 | |
| Llama3.2-11B-Vision-InstructInput=Image2025.06 | 15.84 | 26.55 | |
| Mistral-7B-Instruct-v0.3Input=Text2025.06 | 15.61 | 31.88 | |
| TableLlamaInput=Text2025.06 | 14.3 | 19.35 | |
| DeepSeek-R1-Distill-Qwen-1.5BInput Modality=Text2025.06 | 11.18 | 17.75 | |
| mPLUG-Owl3-7BInput=Image2025.06 | 8.71 | 14.34 | |
| Code LoopBackbone=Qwen3-1.7B2026.03 | 6.91 | 7.65 | |
| LLaVa-v1.5-7BInput=Image2025.06 | 5.51 | 9.37 | |
| Table-LLava-7BInput=Image2025.06 | 4.19 | 7.05 | |
| mPLUG-Owl2-7BInput Modality=Image2025.06 | 3.46 | 7.23 | |
| Table-R1-Zero-7BModel Scale Group=Comparable2025.12 | 0 | — |