Data Analysis on RealHitBench
79.55GPT ScoreDeepSeek-R1
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| DeepSeek-R1Input=Text2025.06 | 79.55 | 42.59 | |
| GPT4o(TreeThinker)Input=Image+Text2025.06 | 79.45 | 37.08 | |
| GPT4o(TreeThinker)Input=Text2025.06 | 77.26 | 37.63 | |
| Doubao-1.5-pro-32kInput Modality=Text2025.06 | 75.21 | 36.01 | |
| Gemini1.5-proInput=Image+Text2025.06 | 74.8 | 36.26 | |
| GPT4oInput=Text2025.06 | 73.37 | 36.36 | |
| GPT4oInput=Image+Text2025.06 | 72.05 | 35.25 | |
| DTRBackbone=DeepSeek-v32026.03 | 70.9 | 38.67 | |
| GPT4o(TreeThinker)Input=Image2025.06 | 70.83 | 34.44 | |
| Gemini1.5-proInput=Text2025.06 | 70.72 | 36.17 | |
| QwQ-32BInput Modality=Text2025.06 | 68.66 | 20.86 | |
| Qwen2.5-72B-InstructInput=Text2025.06 | 68.45 | 35.9 | |
| Gemini1.5-proInput=Image2025.06 | 67.22 | 36.26 | |
| Qwen3-32BModel Scale Group=Larger2025.12 | 66.67 | — | |
| TableGPT-R1-8BModel Scale Group=Comparable2025.12 | 66.53 | — | |
| DeepSeek-V3Model Scale Group=Larger2025.12 | 66.29 | — | |
| GPT4oInput=Image2025.06 | 65.24 | 33.1 | |
| GPT4o2026.03 | 65.24 | 33.1 | |
| QwQ-32BModel Scale Group=Larger2025.12 | 64.99 | — | |
| Qwen3-14BModel Scale Group=Larger2025.12 | 63.03 | — | |
| TableGPT2-7BInput=Text2025.06 | 62.76 | 33.25 | |
| TableGPT2-7B2026.03 | 62.76 | 33.25 | |
| Code LoopBackbone=DeepSeek-v32026.03 | 62.51 | 33.73 | |
| Qwen-PlusModel Scale Group=Larger2025.12 | 62.04 | — | |
| DeepSeek-v32026.03 | 61.4 | 34.76 | |
| Llama-3.1-8BModel Scale Group=Comparable2025.12 | 60.12 | — | |
| Llama3.1-8B-InstructInput=Text2025.06 | 60.12 | 32.25 | |
| Mistral-7B-Instruct-v0.3Input=Text2025.06 | 57.74 | 19.84 | |
| Llama3.2-90B-Vision-InstructInput=Text2025.06 | 57.28 | 31.87 | |
| TableLLM-7B2026.03 | 55.8 | 28.42 | |
| GPT-4oModel Scale Group=Larger2025.12 | 55.54 | — | |
| DeepSeek-R1-Distill-Llama-8BInput Modality=Text2025.06 | 54.55 | 26.98 | |
| DeepSeek-R1-Distill-Qwen-7BInput Modality=Text2025.06 | 54.39 | 25.19 | |
| DTRBackbone=Qwen3-4B2026.03 | 53.88 | 20.15 | |
| Qwen3-8BModel Scale Group=Comparable2025.12 | 53.28 | — | |
| Qwen3-72BModel Scale Group=Larger2025.12 | 53.27 | — | |
| TableGPT2-7BModel Scale Group=Comparable2025.12 | 53.16 | — | |
| Llama3.2-90B-Vision-InstructInput=Image+Text2025.06 | 53.06 | 32.36 | |
| Llama3.2-11B-Vision-InstructInput=Text2025.06 | 52.75 | 30.59 | |
| Llama3.3-70B-InstructInput=Text2025.06 | 52.26 | 27.98 | |
| TableLLMModel Scale Group=Comparable2025.12 | 47.86 | — | |
| TableLLM-Llama3.1-8BInput=Text2025.06 | 47.86 | 27.26 | |
| TableLLM-Qwen2-7BInput=Text2025.06 | 47.3 | 18.86 | |
| StructGPT2026.03 | 42.33 | 19.5 | |
| Llama3.2-11B-Vision-InstructInput=Image2025.06 | 41.32 | 22.75 | |
| Llama3.2-90B-Vision-InstructInput=Image2025.06 | 41.18 | 25.46 | |
| DTRBackbone=Qwen3-1.7B2026.03 | 40.28 | 16.44 | |
| Qwen2.5-7B-InstructInput=Text2025.06 | 40.17 | 24.17 | |
| Llama3.2-11B-Vision-InstructInput=Image+Text2025.06 | 39.98 | 22.43 | |
| Qwen2-VL-7B-InstructInput Modality=Image+Text2025.06 | 39.69 | 25.39 | |
| DeepSeek-R1-Distill-Qwen-1.5BInput Modality=Text2025.06 | 38.02 | 16.24 | |
| Qwen2-VL-7B-InstructInput Modality=Image2025.06 | 37.41 | 24.3 | |
| Table-R1-Zero-7BModel Scale Group=Comparable2025.12 | 36.24 | — | |
| mPLUG-Owl3-7BInput=Image2025.06 | 30.71 | 18.53 | |
| LLaVa-v1.5-7BInput=Image2025.06 | 26.6 | 17.74 | |
| Code LoopBackbone=Qwen3-4B2026.03 | 25.83 | 18.17 | |
| mPLUG-Owl2-7BInput Modality=Image2025.06 | 25.46 | 12.13 | |
| TableLlamaInput=Text2025.06 | 24.73 | 7.89 | |
| Table-LLava-7BInput=Image2025.06 | 22.46 | 8.72 | |
| Code LoopBackbone=Qwen3-1.7B2026.03 | 17.97 | 17.29 |