Automatic Report Evaluation on USAFact
3.85ReadabilityDA.
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| DA.Base Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL-72B-Instruct, Evaluator=GPT-4.1, Measurement=Average Rank2026.01 | 3.85 | 3.95 | 4 | 4 | 3.45 | 3.2 | |
| DA.Base Model=Qwen3-VL-32B-Instruct, Level=chart2026.01 | 3.4 | 3.65 | — | — | — | — | |
| DirectBase Model=Qwen3-VL-32B-Instruct, Level=chart2026.01 | 3.3 | 2.95 | — | — | — | — | |
| DirectBase Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL-72B-Instruct, Evaluator=GPT-4.1, Measurement=Average Rank2026.01 | 2.85 | 2.8 | 2.1 | 2.8 | 2.9 | 2.25 | |
| DN.Base Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL-72B-Instruct, Evaluator=GPT-4.1, Measurement=Average Rank2026.01 | 1.9 | 1.65 | 2.05 | 2.1 | 2.25 | 2.65 | |
| DN.Base Model=Qwen3-VL-32B-Instruct, Level=chart2026.01 | 1.8 | 1.3 | — | — | — | — | |
| EvidFuseBase Model=Qwen3-VL-32B-Instruct, Level=chart2026.01 | 1.5 | 2.1 | — | — | — | — | |
| EvidFuseBase Models=Qwen3-235B-A22B-Instruct-2507 & Qwen2.5-VL-72B-Instruct, Evaluator=GPT-4.1, Measurement=Average Rank2026.01 | 1.4 | 1.6 | 1.85 | 1.1 | 1.4 | 1.7 | |
| DA.Base Model=Qwen3-VL-32B-Instruct, Level=chapter2026.01 | — | — | 4 | 4 | — | — | |
| DA.Base Model=Qwen3-VL-32B-Instruct, Level=report2026.01 | — | — | — | — | 3.25 | 3.15 | |
| DirectBase Model=Qwen3-VL-32B-Instruct, Level=chapter2026.01 | — | — | 2.65 | 2.5 | — | — | |
| DirectBase Model=Qwen3-VL-32B-Instruct, Level=report2026.01 | — | — | — | — | 3.5 | 2.5 | |
| DN.Base Model=Qwen3-VL-32B-Instruct, Level=chapter2026.01 | — | — | 2.2 | 2.4 | — | — | |
| DN.Base Model=Qwen3-VL-32B-Instruct, Level=report2026.01 | — | — | — | — | 1.75 | 2.5 | |
| EvidFuseBase Model=Qwen3-VL-32B-Instruct, Level=chapter2026.01 | — | — | 1.15 | 1.1 | — | — | |
| EvidFuseBase Model=Qwen3-VL-32B-Instruct, Level=report2026.01 | — | — | — | — | 1.5 | 1.65 |