Automatic Report Evaluation on Tableau
3.65Readability ScoreDA.
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| DA.Base Model=Qwen3-VL-32B-Instruct, Level=chart2026.01 | 3.65 | 3.85 | — | — | — | — | |
| DA.Base Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL-72B-Instruct, Evaluator=GPT-4.1, Measurement=Average Rank2026.01 | 3.6 | 3.9 | 4 | 4 | 3.75 | 3.05 | |
| DeepAnalyzeLevel=chart, Base Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL 72B-Instruct2026.01 | 3.4 | 3.53 | — | — | — | — | |
| DirectBase Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL-72B-Instruct, Evaluator=GPT-4.1, Measurement=Average Rank2026.01 | 3.1 | 2.95 | 1.95 | 2.15 | 2.9 | 2.55 | |
| DirectLevel=chart, Base Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL 72B-Instruct2026.01 | 3.1 | 2.93 | — | — | — | — | |
| DirectBase Model=Qwen3-VL-32B-Instruct, Level=chart2026.01 | 3.05 | 2.8 | — | — | — | — | |
| DataNarrativeLevel=chart, Base Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL 72B-Instruct2026.01 | 2.27 | 2.23 | — | — | — | — | |
| EvidFuseBase Model=Qwen3-VL-32B-Instruct, Level=chart2026.01 | 1.85 | 1.9 | — | — | — | — | |
| DN.Base Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL-72B-Instruct, Evaluator=GPT-4.1, Measurement=Average Rank2026.01 | 1.85 | 1.35 | 2.35 | 2.3 | 1.75 | 2.6 | |
| DN.Base Model=Qwen3-VL-32B-Instruct, Level=chart2026.01 | 1.45 | 1.45 | — | — | — | — | |
| EvidFuseBase Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL-72B-Instruct, Evaluator=GPT-4.1, Measurement=Average Rank2026.01 | 1.45 | 1.8 | 1.7 | 1.55 | 1.6 | 1.8 | |
| EvidFuseLevel=chart, Base Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL 72B-Instruct2026.01 | 1.23 | 1.3 | — | — | — | — | |
| DA.Base Model=Qwen3-VL-32B-Instruct, Level=chapter2026.01 | — | — | 4 | 4 | — | — | |
| DA.Base Model=Qwen3-VL-32B-Instruct, Level=report2026.01 | — | — | — | — | 3.55 | 3.1 | |
| DataNarrativeLevel=chapter, Base Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL 72B-Instruct2026.01 | — | — | 2.7 | 2.3 | — | — | |
| DataNarrativeLevel=report, Base Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL 72B-Instruct2026.01 | — | — | — | — | 2.1 | 2.7 | |
| DeepAnalyzeLevel=chapter, Base Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL 72B-Instruct2026.01 | — | — | 4 | 3.37 | — | — | |
| DeepAnalyzeLevel=report, Base Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL 72B-Instruct2026.01 | — | — | — | — | 3.07 | 3.2 | |
| DirectBase Model=Qwen3-VL-32B-Instruct, Level=chapter2026.01 | — | — | 2.05 | 2.35 | — | — | |
| DirectBase Model=Qwen3-VL-32B-Instruct, Level=report2026.01 | — | — | — | — | 3.05 | 2.7 | |
| DirectLevel=chapter, Base Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL 72B-Instruct2026.01 | — | — | 2.23 | 3.2 | — | — | |
| DirectLevel=report, Base Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL 72B-Instruct2026.01 | — | — | — | — | 3.6 | 2.7 | |
| DN.Base Model=Qwen3-VL-32B-Instruct, Level=chapter2026.01 | — | — | 2.7 | 2.45 | — | — | |
| DN.Base Model=Qwen3-VL-32B-Instruct, Level=report2026.01 | — | — | — | — | 1.95 | 2.5 | |
| EvidFuseBase Model=Qwen3-VL-32B-Instruct, Level=chapter2026.01 | — | — | 1.25 | 1.2 | — | — | |
| EvidFuseBase Model=Qwen3-VL-32B-Instruct, Level=report2026.01 | — | — | — | — | 1.45 | 1.7 | |
| EvidFuseLevel=chapter, Base Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL 72B-Instruct2026.01 | — | — | 1.07 | 1.13 | — | — | |
| EvidFuseLevel=report, Base Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL 72B-Instruct2026.01 | — | — | — | — | 1.23 | 1.4 |