Chart Generation on RealHitBench
100ECRDTR
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| DTRBackbone=DeepSeek-v32026.03 | 100 | 52.69 | |
| GPT4o(TreeThinker)Input=Text2025.06 | 67.76 | 39.47 | |
| TableGPT2-7BInput=Text2025.06 | 67.53 | 32.47 | |
| TableGPT2-7B2026.03 | 67.53 | 32.47 | |
| GPT4o(TreeThinker)Input=Image2025.06 | 67.32 | 19.61 | |
| GPT4o(TreeThinker)Input=Image+Text2025.06 | 65.13 | 33.55 | |
| TableGPT-R1-8BModel Scale Group=Comparable2025.12 | 55.84 | — | |
| Llama3.3-70B-InstructInput=Text2025.06 | 50.65 | 24.03 | |
| Qwen2.5-7B-InstructInput=Text2025.06 | 48.7 | 15.58 | |
| Qwen-PlusModel Scale Group=Larger2025.12 | 48.05 | — | |
| TableGPT2-7BModel Scale Group=Comparable2025.12 | 44.16 | — | |
| TableLLM-7B2026.03 | 42.15 | 18.2 | |
| GPT4oInput=Text2025.06 | 40.26 | 20.13 | |
| Llama3.2-90B-Vision-InstructInput=Image+Text2025.06 | 39.61 | 13.64 | |
| DeepSeek-R1-Distill-Qwen-7BInput Modality=Text2025.06 | 39.61 | 6.49 | |
| Code LoopBackbone=DeepSeek-v32026.03 | 39.6 | 20.78 | |
| Gemini1.5-proInput=Image2025.06 | 38.96 | 7.79 | |
| Llama3.2-11B-Vision-InstructInput=Image+Text2025.06 | 35.71 | 6.49 | |
| GPT-4oModel Scale Group=Larger2025.12 | 34.42 | — | |
| DeepSeek-R1Input=Text2025.06 | 32.14 | 7.14 | |
| QwQ-32BInput Modality=Text2025.06 | 31.25 | 12.5 | |
| GPT4oInput=Image+Text2025.06 | 30.52 | 14.29 | |
| DTRBackbone=Qwen3-4B2026.03 | 30.32 | 8.39 | |
| Qwen2.5-72B-InstructInput=Text2025.06 | 27.27 | 14.29 | |
| Doubao-1.5-pro-32kInput Modality=Text2025.06 | 25.97 | 14.94 | |
| GPT4oInput=Image2025.06 | 25.32 | 10.39 | |
| Gemini1.5-proInput=Text2025.06 | 25.32 | 9.74 | |
| Qwen2-VL-7B-InstructInput Modality=Image+Text2025.06 | 25.32 | 9.09 | |
| GPT4o2026.03 | 25.32 | 10.39 | |
| Qwen3-32BModel Scale Group=Larger2025.12 | 25 | — | |
| DeepSeek-v32026.03 | 24.68 | 9.09 | |
| Qwen3-8BModel Scale Group=Comparable2025.12 | 24.67 | — | |
| Gemini1.5-proInput=Image+Text2025.06 | 24.03 | 9.09 | |
| Qwen3-14BModel Scale Group=Larger2025.12 | 23.38 | — | |
| TableLLMModel Scale Group=Comparable2025.12 | 22.73 | — | |
| TableLLM-Llama3.1-8BInput=Text2025.06 | 22.73 | 6.49 | |
| Llama3.2-90B-Vision-InstructInput=Text2025.06 | 22.73 | 7.79 | |
| DTRBackbone=Qwen3-1.7B2026.03 | 21.94 | 5.16 | |
| Qwen3-72BModel Scale Group=Larger2025.12 | 20.78 | — | |
| Llama3.2-11B-Vision-InstructInput=Image2025.06 | 20.78 | 1.95 | |
| QwQ-32BModel Scale Group=Larger2025.12 | 20.13 | — | |
| DeepSeek-V3Model Scale Group=Larger2025.12 | 18.18 | — | |
| Mistral-7B-Instruct-v0.3Input=Text2025.06 | 18.18 | 7.14 | |
| Llama3.2-11B-Vision-InstructInput=Text2025.06 | 17.53 | 4.55 | |
| Llama3.2-90B-Vision-InstructInput=Image2025.06 | 16.23 | 5.19 | |
| Qwen2-VL-7B-InstructInput Modality=Image2025.06 | 16.23 | 5.84 | |
| Table-R1-Zero-7BModel Scale Group=Comparable2025.12 | 16 | — | |
| DeepSeek-R1-Distill-Qwen-1.5BInput Modality=Text2025.06 | 14.94 | 0 | |
| Code LoopBackbone=Qwen3-4B2026.03 | 14.3 | 3.25 | |
| Llama-3.1-8BModel Scale Group=Comparable2025.12 | 13.64 | — | |
| Llama3.1-8B-InstructInput=Text2025.06 | 13.64 | 4.55 | |
| DeepSeek-R1-Distill-Llama-8BInput Modality=Text2025.06 | 12.99 | 5.19 | |
| StructGPT2026.03 | 12.44 | 5.12 | |
| LLaVa-v1.5-7BInput=Image2025.06 | 11.04 | 0 | |
| TableLLM-Qwen2-7BInput=Text2025.06 | 10.39 | 4.55 | |
| mPLUG-Owl3-7BInput=Image2025.06 | 6.49 | 3.25 | |
| mPLUG-Owl2-7BInput Modality=Image2025.06 | 1.53 | 0 | |
| Table-LLava-7BInput=Image2025.06 | 1.3 | 0 | |
| Code LoopBackbone=Qwen3-1.7B2026.03 | 0.6 | 0 | |
| TableLlamaInput=Text2025.06 | 0 | 0 |