Numerical Reasoning on RealHitBench
70.31Exact Match (EM)DeepSeek-R1
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| DeepSeek-R1Prompting Strategy=Direct2026.04 | 70.31 | 72.54 | |
| SpreadsheetAgentModel Backbone=Llama3.3-70B-Instruct, Prompting Strategy=Explicit Reasoning2026.04 | 58.37 | 64.44 | |
| GPT-OSS-120B w/ SpreadsheetAgentTool=Python2026.04 | 58.24 | 62.58 | |
| TreeThinkerModel Backbone=Llama3.3-70B-Instruct, Prompting Strategy=Explicit Reasoning2026.04 | 57.98 | 64.92 | |
| DTRBackbone=DeepSeek-v32026.03 | 55.51 | 61.98 | |
| QwQ-32BModel Scale Group=Larger2025.12 | 55.38 | — | |
| GPT-OSS-120BTool=Python2026.04 | 55.38 | 59.16 | |
| Qwen3-Coder-480B w/ SpreadsheetAgentTool=Python2026.04 | 55.25 | 65.74 | |
| GPT-OSS-120B w/ TreeThinkerTool=Python2026.04 | 54.6 | 59.06 | |
| Qwen3-Coder-480BTool=Python2026.04 | 54.6 | 64.97 | |
| Llama3.3-70B-InstructPrompting Strategy=Explicit Reasoning2026.04 | 54.47 | 61.6 | |
| SpreadsheetAgentModel Backbone=Qwen2.5-72B-Instruct, Prompting Strategy=Explicit Reasoning2026.04 | 53.05 | 62.31 | |
| DeepSeek-V3Model Scale Group=Larger2025.12 | 52.53 | — | |
| Qwen2.5-72B-InstructPrompting Strategy=Explicit Reasoning2026.04 | 52.4 | 61.33 | |
| TreeThinkerModel Backbone=Qwen2.5-72B-Instruct, Prompting Strategy=Explicit Reasoning2026.04 | 50.97 | 60.33 | |
| TableGPT-R1-8BModel Scale Group=Comparable2025.12 | 49.03 | — | |
| Qwen3-235B w/ SpreadsheetAgentTool=Python2026.04 | 47.47 | 54.25 | |
| Qwen3-32BModel Scale Group=Larger2025.12 | 47.34 | — | |
| Qwen3-Coder-480B w/ TreeThinkerTool=Python2026.04 | 47.21 | 58.05 | |
| DeepSeek-v32026.03 | 47.05 | 50.61 | |
| Qwen3-72BModel Scale Group=Larger2025.12 | 46.95 | — | |
| Doubao-1.5-pro-32kInput Modality=Text2025.06 | 45.4 | 53.77 | |
| GPT-OSS-20B w/ SpreadsheetAgentTool=Python2026.04 | 44.62 | 48.92 | |
| Qwen3-235B w/ TreeThinkerTool=Python2026.04 | 44.1 | 50.06 | |
| Qwen3-14BModel Scale Group=Larger2025.12 | 43.7 | — | |
| Code LoopBackbone=DeepSeek-v32026.03 | 42.93 | 49.68 | |
| Qwen3-235BTool=Python2026.04 | 42.54 | 48.94 | |
| GPT-OSS-20B w/ TreeThinkerTool=Python2026.04 | 39.82 | 44.42 | |
| Qwen3-8BModel Scale Group=Comparable2025.12 | 39.43 | — | |
| GPT-4oModel Scale Group=Larger2025.12 | 38.91 | — | |
| GPT4oPrompting Strategy=Direct2026.04 | 38.65 | 50.12 | |
| Llama3.3-70B-InstructPrompting Strategy=Direct2026.04 | 36.58 | 48.99 | |
| Qwen3-30B w/ SpreadsheetAgentTool=Python2026.04 | 35.8 | 44.36 | |
| Gemini1.5-proPrompting Strategy=Direct2026.04 | 35.54 | 43.74 | |
| GPT-OSS-20BTool=Python2026.04 | 34.63 | 39.03 | |
| SpreadsheetAgentModel Backbone=Qwen2.5-7B-Instruct, Prompting Strategy=Explicit Reasoning2026.04 | 32.3 | 40.44 | |
| Qwen3-30B w/ TreeThinkerTool=Python2026.04 | 31.26 | 39.77 | |
| Qwen-PlusModel Scale Group=Larger2025.12 | 31.25 | — | |
| TreeThinkerModel Backbone=Qwen2.5-7B-Instruct, Prompting Strategy=Explicit Reasoning2026.04 | 31.13 | 39.14 | |
| Qwen2.5-7B-InstructPrompting Strategy=Explicit Reasoning2026.04 | 29.83 | 38.29 | |
| TableGPT2-7B2026.03 | 29.31 | 39.81 | |
| Qwen3-30BTool=Python2026.04 | 28.66 | 35.76 | |
| GPT4o2026.03 | 27.63 | 36.68 | |
| Qwen2.5-72B-InstructPrompting Strategy=Direct2026.04 | 26.98 | 39.23 | |
| TableGPT2-7BModel Scale Group=Comparable2025.12 | 24.9 | — | |
| TableLLM-7B2026.03 | 22.4 | 31.65 | |
| DTRBackbone=Qwen3-4B2026.03 | 22.05 | 30.13 | |
| QwQ-32BInput Modality=Text2025.06 | 17.94 | 33.57 | |
| DeepSeek-R1-Distill-Qwen-7BInput Modality=Text2025.06 | 17.64 | 24.96 | |
| SpreadsheetAgentModel Backbone=Llama3.1-8B-Instruct, Prompting Strategy=Explicit Reasoning2026.04 | 15.69 | 28.36 | |
| DeepSeek-R1-Distill-Llama-8BInput Modality=Text2025.06 | 15.18 | 24.64 | |
| StructGPT2026.03 | 14.55 | 20.8 | |
| Llama-3.1-8BModel Scale Group=Comparable2025.12 | 14.53 | — | |
| Llama3.1-8B-InstructPrompting Strategy=Direct2026.04 | 14.53 | 27.21 | |
| DTRBackbone=Qwen3-1.7B2026.03 | 13.75 | 19.01 | |
| TableLLMModel Scale Group=Comparable2025.12 | 13.36 | — | |
| TreeThinkerModel Backbone=Llama3.1-8B-Instruct, Prompting Strategy=Explicit Reasoning2026.04 | 13.36 | 24.76 | |
| Qwen2-VL-7B-InstructInput Modality=Image+Text2025.06 | 11.54 | 24.97 | |
| Llama3.1-8B-InstructPrompting Strategy=Explicit Reasoning2026.04 | 10.38 | 22.58 | |
| Code LoopBackbone=Qwen3-4B2026.03 | 10.25 | 13 | |
| Qwen2-VL-7B-InstructInput Modality=Image2025.06 | 9.34 | 22.88 | |
| DeepSeek-R1-Distill-Qwen-1.5BInput Modality=Text2025.06 | 7.13 | 15.74 | |
| Qwen2.5-7B-InstructPrompting Strategy=Direct2026.04 | 5.32 | 19.75 | |
| Code LoopBackbone=Qwen3-1.7B2026.03 | 4.02 | 5.75 | |
| mPLUG-Owl2-7BInput Modality=Image2025.06 | 1.15 | 4.72 | |
| Table-R1-Zero-7BModel Scale Group=Comparable2025.12 | 0 | — |