Automatic Report Evaluation on OurWorld InData
3.7ReadabilityDA.
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| DA.Base Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL-72B-Instruct, Evaluator=GPT-4.1, Measurement=Average Rank2026.01 | 3.7 | 3.9 | 4 | 4 | 3.5 | 3.45 | |
| DA.Base Model=Qwen3-VL-32B-Instruct, Level=chart2026.01 | 3.4 | 3.6 | — | — | — | — | |
| DirectBase Model=Qwen3-VL-32B-Instruct, Level=chart2026.01 | 3.3 | 3.25 | — | — | — | — | |
| DirectLevel=chart, Base Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL 72B-Instruct2026.01 | 3.23 | 3.07 | — | — | — | — | |
| DeepAnalyzeLevel=chart, Base Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL 72B-Instruct2026.01 | 3.17 | 3.5 | — | — | — | — | |
| DirectBase Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL-72B-Instruct, Evaluator=GPT-4.1, Measurement=Average Rank2026.01 | 2.85 | 2.85 | 1.95 | 2.2 | 2.75 | 2.05 | |
| DataNarrativeLevel=chart, Base Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL 72B-Instruct2026.01 | 2.27 | 2.2 | — | — | — | — | |
| DN.Base Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL-72B-Instruct, Evaluator=GPT-4.1, Measurement=Average Rank2026.01 | 2.12 | 1.75 | 2.35 | 2.4 | 2.25 | 2.85 | |
| DN.Base Model=Qwen3-VL-32B-Instruct, Level=chart2026.01 | 1.8 | 1.4 | — | — | — | — | |
| EvidFuseBase Model=Qwen3-VL-32B-Instruct, Level=chart2026.01 | 1.5 | 1.7 | — | — | — | — | |
| EvidFuseLevel=chart, Base Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL 72B-Instruct2026.01 | 1.33 | 1.23 | — | — | — | — | |
| EvidFuseBase Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL-72B-Instruct, Evaluator=GPT-4.1, Measurement=Average Rank2026.01 | 1.3 | 1.5 | 1.7 | 1.4 | 1.5 | 1.65 | |
| DA.Base Model=Qwen3-VL-32B-Instruct, Level=chapter2026.01 | — | — | 4 | 4 | — | — | |
| DA.Base Model=Qwen3-VL-32B-Instruct, Level=report2026.01 | — | — | — | — | 3.35 | 2.7 | |
| DataNarrativeLevel=chapter, Base Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL 72B-Instruct2026.01 | — | — | 2.6 | 2.6 | — | — | |
| DataNarrativeLevel=report, Base Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL 72B-Instruct2026.01 | — | — | — | — | 2.23 | 2.53 | |
| DeepAnalyzeLevel=chapter, Base Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL 72B-Instruct2026.01 | — | — | 4 | 3.57 | — | — | |
| DeepAnalyzeLevel=report, Base Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL 72B-Instruct2026.01 | — | — | — | — | 3.2 | 3.23 | |
| DirectBase Model=Qwen3-VL-32B-Instruct, Level=chapter2026.01 | — | — | 2.2 | 2.55 | — | — | |
| DirectBase Model=Qwen3-VL-32B-Instruct, Level=report2026.01 | — | — | — | — | 3.1 | 2.55 | |
| DirectLevel=chapter, Base Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL 72B-Instruct2026.01 | — | — | 2.4 | 2.8 | — | — | |
| DirectLevel=report, Base Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL 72B-Instruct2026.01 | — | — | — | — | 3.33 | 2.77 | |
| DN.Base Model=Qwen3-VL-32B-Instruct, Level=chapter2026.01 | — | — | 2.75 | 2.2 | — | — | |
| DN.Base Model=Qwen3-VL-32B-Instruct, Level=report2026.01 | — | — | — | — | 2.2 | 2.85 | |
| EvidFuseBase Model=Qwen3-VL-32B-Instruct, Level=chapter2026.01 | — | — | 1.05 | 1.25 | — | — | |
| EvidFuseBase Model=Qwen3-VL-32B-Instruct, Level=report2026.01 | — | — | — | — | 1.35 | 1.5 | |
| EvidFuseLevel=chapter, Base Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL 72B-Instruct2026.01 | — | — | 1 | 1.03 | — | — | |
| EvidFuseLevel=report, Base Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL 72B-Instruct2026.01 | — | — | — | — | 1.23 | 1.47 |