Multimodal Document Question Answering on DocBench (test)
52.4Accuracy (Academic)BayesRAG
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| BayesRAGResponse Generator=GPT-4o-mini, Parsing=MinerU2026.01 | 52.4 | 37.5 | 56 | 52.3 | 66.8 | 51.2 | |
| RAGAnythingResponse Generator=GPT-4o-mini, Paradigm=Knowledge-graph-based2026.01 | 51.4 | 42.3 | 43.2 | 50.2 | 57.5 | 48.7 | |
| RAGFlowResponse Generator=GPT-4o-mini, Paradigm=Modular RAG2026.01 | 45.2 | 34 | 41.8 | 48.1 | 23.8 | 39 | |
| GPT-4o-miniInput=Page-level screenshots (max 50), Direct evaluation=true2026.01 | 42.5 | 36.8 | 53.1 | 51.4 | 51.1 | 44.1 | |
| ViDoRAGResponse Generator=GPT-4o-mini, Agentic Workflow=Explore, Summarize, Reflect2026.01 | 29 | 69.1 | 35.8 | 39.7 | 37.2 | 43.5 | |
| VisRAGResponse Generator=GPT-4o-mini, Parsing=Visual-centric (no OCR)2026.01 | 21.9 | 28.2 | 20.3 | 29.6 | 11.6 | 25.3 |