Structure Comprehending on RealHitBench
82.71Exact Match (EM)DeepSeek-R1
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| DeepSeek-R1Input=Text2025.06 | 82.71 | 84.62 | |
| DeepSeek-R1Prompting Strategy=Direct2026.04 | 82.71 | 84.62 | |
| Qwen3-Coder-480B w/ SpreadsheetAgentTool=Python2026.04 | 76.26 | 82.71 | |
| QwQ-32BModel Scale Group=Larger2025.12 | 76.08 | — | |
| TreeThinkerModel Backbone=Llama3.3-70B-Instruct, Prompting Strategy=Explicit Reasoning2026.04 | 74.24 | 79.49 | |
| SpreadsheetAgentModel Backbone=Llama3.3-70B-Instruct, Prompting Strategy=Explicit Reasoning2026.04 | 74.24 | 79.36 | |
| Llama3.3-70B-InstructPrompting Strategy=Explicit Reasoning2026.04 | 73.99 | 78.61 | |
| Qwen2.5-72B-InstructPrompting Strategy=Explicit Reasoning2026.04 | 73.48 | 80.54 | |
| Qwen3-14BModel Scale Group=Larger2025.12 | 73.02 | — | |
| Qwen3-Coder-480BTool=Python2026.04 | 72.98 | 80.01 | |
| TreeThinkerModel Backbone=Qwen2.5-72B-Instruct, Prompting Strategy=Explicit Reasoning2026.04 | 72.73 | 79.45 | |
| Qwen3-32BModel Scale Group=Larger2025.12 | 71.76 | — | |
| DeepSeek-V3Model Scale Group=Larger2025.12 | 71.25 | — | |
| SpreadsheetAgentModel Backbone=Qwen2.5-72B-Instruct, Prompting Strategy=Explicit Reasoning2026.04 | 70.71 | 78.42 | |
| Qwen3-72BModel Scale Group=Larger2025.12 | 69.47 | — | |
| GPT4o(TreeThinker)Input=Image+Text2025.06 | 66.31 | 77.42 | |
| GPT4oInput=Image+Text2025.06 | 65.91 | 74.49 | |
| Doubao-1.5-pro-32kInput Modality=Text2025.06 | 65.9 | 72.45 | |
| GPT4o(TreeThinker)Input=Text2025.06 | 64.4 | 75.67 | |
| Qwen3-8BModel Scale Group=Comparable2025.12 | 64.12 | — | |
| TableGPT-R1-8BModel Scale Group=Comparable2025.12 | 64.12 | — | |
| Qwen3-235B w/ SpreadsheetAgentTool=Python2026.04 | 63.89 | 69.43 | |
| Gemini1.5-proInput=Text2025.06 | 63.64 | 69.71 | |
| Gemini1.5-proPrompting Strategy=Direct2026.04 | 63.64 | 69.71 | |
| GPT4oInput=Text2025.06 | 63.04 | 71.14 | |
| GPT4oPrompting Strategy=Direct2026.04 | 63.04 | 71.14 | |
| Qwen-PlusModel Scale Group=Larger2025.12 | 62.85 | — | |
| GPT-OSS-120B w/ SpreadsheetAgentTool=Python2026.04 | 62.37 | 66.01 | |
| Qwen3-235BTool=Python2026.04 | 61.87 | 67.85 | |
| GPT-4oModel Scale Group=Larger2025.12 | 61.83 | — | |
| Llama3.2-90B-Vision-InstructInput=Image+Text2025.06 | 59.8 | 70.23 | |
| Qwen3-235B w/ TreeThinkerTool=Python2026.04 | 59.09 | 63.28 | |
| Llama3.2-90B-Vision-InstructInput=Text2025.06 | 58.52 | 69.57 | |
| Qwen3-Coder-480B w/ TreeThinkerTool=Python2026.04 | 58.33 | 65.77 | |
| Gemini1.5-proInput=Image+Text2025.06 | 56.74 | 65.8 | |
| DTRBackbone=DeepSeek-v32026.03 | 56.57 | 77.95 | |
| Llama3.3-70B-InstructInput=Text2025.06 | 55.81 | 68.93 | |
| Llama3.3-70B-InstructPrompting Strategy=Direct2026.04 | 55.81 | 68.93 | |
| Qwen2.5-72B-InstructInput=Text2025.06 | 54.55 | 68.34 | |
| Qwen2.5-72B-InstructPrompting Strategy=Direct2026.04 | 54.55 | 68.34 | |
| GPT-OSS-120BTool=Python2026.04 | 54.04 | 57.19 | |
| TableLLMModel Scale Group=Comparable2025.12 | 53.28 | — | |
| TableLLM-Llama3.1-8BInput=Text2025.06 | 53.28 | 58.72 | |
| GPT-OSS-120B w/ TreeThinkerTool=Python2026.04 | 53.03 | 57.98 | |
| Code LoopBackbone=Qwen3-1.7B2026.03 | 49.24 | 53.08 | |
| GPT4o(TreeThinker)Input=Image2025.06 | 49.21 | 58.32 | |
| SpreadsheetAgentModel Backbone=Qwen2.5-7B-Instruct, Prompting Strategy=Explicit Reasoning2026.04 | 48.74 | 56.61 | |
| Qwen2-VL-7B-InstructInput Modality=Image+Text2025.06 | 48.6 | 60.42 | |
| TableGPT2-7BInput=Text2025.06 | 48.23 | 56.68 | |
| TableGPT2-7B2026.03 | 48.23 | 56.68 | |
| Qwen2.5-7B-InstructPrompting Strategy=Explicit Reasoning2026.04 | 45.96 | 55.45 | |
| Qwen2-VL-7B-InstructInput Modality=Image2025.06 | 45.29 | 55.99 | |
| TreeThinkerModel Backbone=Qwen2.5-7B-Instruct, Prompting Strategy=Explicit Reasoning2026.04 | 44.95 | 53.63 | |
| Code LoopBackbone=DeepSeek-v32026.03 | 44.19 | 51.95 | |
| DeepSeek-v32026.03 | 43.31 | 74.63 | |
| GPT4oInput=Image2025.06 | 42.68 | 52.89 | |
| GPT4o2026.03 | 42.68 | 52.89 | |
| Gemini1.5-proInput=Image2025.06 | 41.52 | 50.23 | |
| mPLUG-Owl3-7BInput=Image2025.06 | 41.24 | 48.86 | |
| TableLLM-7B2026.03 | 41.1 | 49.34 | |
| Qwen3-30B w/ SpreadsheetAgentTool=Python2026.04 | 40.15 | 47.28 | |
| GPT-OSS-20B w/ SpreadsheetAgentTool=Python2026.04 | 39.65 | 44.74 | |
| Qwen3-30BTool=Python2026.04 | 37.88 | 44.26 | |
| TableLlamaInput=Text2025.06 | 36.31 | 42.91 | |
| TableLLM-Qwen2-7BInput=Text2025.06 | 36.09 | 44.75 | |
| Llama-3.1-8BModel Scale Group=Comparable2025.12 | 35.9 | — | |
| Llama3.1-8B-InstructInput=Text2025.06 | 35.9 | 50.8 | |
| Llama3.1-8B-InstructPrompting Strategy=Direct2026.04 | 35.9 | 50.8 | |
| TableGPT2-7BModel Scale Group=Comparable2025.12 | 34.86 | — | |
| Qwen3-30B w/ TreeThinkerTool=Python2026.04 | 34.09 | 41.65 | |
| GPT-OSS-20BTool=Python2026.04 | 33.84 | 40.46 | |
| GPT-OSS-20B w/ TreeThinkerTool=Python2026.04 | 33.59 | 39.53 | |
| Llama3.2-90B-Vision-InstructInput=Image2025.06 | 33.33 | 46.4 | |
| Code LoopBackbone=Qwen3-4B2026.03 | 32.85 | 30.3 | |
| DTRBackbone=Qwen3-4B2026.03 | 32.6 | 43.32 | |
| StructGPT2026.03 | 30.25 | 38.6 | |
| Table-R1-Zero-7BModel Scale Group=Comparable2025.12 | 28.5 | — | |
| Llama3.2-11B-Vision-InstructInput=Image+Text2025.06 | 27.99 | 43.85 | |
| Llama3.1-8B-InstructPrompting Strategy=Explicit Reasoning2026.04 | 25.51 | 39.51 | |
| SpreadsheetAgentModel Backbone=Llama3.1-8B-Instruct, Prompting Strategy=Explicit Reasoning2026.04 | 25.25 | 40.33 | |
| Qwen2.5-7B-InstructInput=Text2025.06 | 23.48 | 44.81 | |
| Qwen2.5-7B-InstructPrompting Strategy=Direct2026.04 | 23.48 | 44.81 | |
| Llama3.2-11B-Vision-InstructInput=Text2025.06 | 23.39 | 43.22 | |
| TreeThinkerModel Backbone=Llama3.1-8B-Instruct, Prompting Strategy=Explicit Reasoning2026.04 | 21.72 | 38.78 | |
| Mistral-7B-Instruct-v0.3Input=Text2025.06 | 21.21 | 49.62 | |
| Llama3.2-11B-Vision-InstructInput=Image2025.06 | 20.87 | 32.57 | |
| DeepSeek-R1-Distill-Qwen-7BInput Modality=Text2025.06 | 19.08 | 28.1 | |
| DTRBackbone=Qwen3-1.7B2026.03 | 17.17 | 27.88 | |
| DeepSeek-R1-Distill-Llama-8BInput Modality=Text2025.06 | 14.76 | 26.66 | |
| LLaVa-v1.5-7BInput=Image2025.06 | 13.23 | 30.51 | |
| QwQ-32BInput Modality=Text2025.06 | 12.99 | 37.96 | |
| DeepSeek-R1-Distill-Qwen-1.5BInput Modality=Text2025.06 | 9.92 | 19.93 | |
| mPLUG-Owl2-7BInput Modality=Image2025.06 | 9.14 | 14.21 | |
| Table-LLava-7BInput=Image2025.06 | 7.38 | 11.83 |