Document Parsing Quality Assessment on DOCRcaseBench Table 1.0 (test)
86.41Case F1DOCR-Inspector-7B
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| DOCR-Inspector-7BCategory=Ours2025.12 | 86.41 | 63.09 | 62.11 | 62.95 | |
| Qwen2.5-VL-72B-InstructCoT=false, Category=Open-source Non-Reasoning Models2025.12 | 83.51 | 40.91 | 33.94 | 31.03 | |
| Qwen3-VL-235B-A22B-ThinkingParadigm=Thinking, Category=Reasoning Models2025.12 | 83.13 | 39.12 | 28.57 | 25.49 | |
| Gemini 2.5 FlashCoT=false, Category=Proprietary Non-Reasoning Models2025.12 | 82.21 | 41.94 | 25.97 | 21.29 | |
| Gemini 2.5 Pro ThinkingParadigm=Thinking, Category=Reasoning Models2025.12 | 82.01 | 43.6 | 32.93 | 29.63 | |
| GPT-4oCoT=true, Category=Proprietary Non-Reasoning Models2025.12 | 81.23 | 34.23 | 29.64 | 28.17 | |
| Gemini 2.5 FlashCoT=true, Category=Proprietary Non-Reasoning Models2025.12 | 81.16 | 42.36 | 24.1 | 19.25 | |
| Qwen2.5-VL-72B-InstructCoT=true, Category=Open-source Non-Reasoning Models2025.12 | 76.82 | 40.7 | 31.77 | 28.43 | |
| GPT-4oCoT=false, Category=Proprietary Non-Reasoning Models2025.12 | 73.69 | 29.89 | 26.36 | 25.03 | |
| Qwen2.5-VL-7B-InstructCoT=false, Category=Open-source Non-Reasoning Models2025.12 | 48.8 | 19.42 | 19.42 | 19.42 | |
| Qwen2.5-VL-7B-InstructCoT=true, Category=Open-source Non-Reasoning Models2025.12 | 43.48 | 21.56 | 21.72 | 22.11 |