Document Parsing Quality Assessment on DOCRcaseBench Equation 1.0 (test)
85.42Case F1DOCR-Inspector-7B
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| DOCR-Inspector-7BCategory=Ours2025.12 | 85.42 | 74.39 | 73.81 | 74.48 | |
| Gemini 2.5 FlashCoT=true, Category=Proprietary Non-Reasoning Models2025.12 | 80.94 | 50.61 | 46.17 | 44.63 | |
| Gemini 2.5 FlashCoT=false, Category=Proprietary Non-Reasoning Models2025.12 | 80.46 | 53.73 | 48.17 | 45.96 | |
| GPT-4oCoT=true, Category=Proprietary Non-Reasoning Models2025.12 | 79.38 | 46.44 | 45.45 | 45.4 | |
| GPT-4oCoT=false, Category=Proprietary Non-Reasoning Models2025.12 | 79.2 | 49.31 | 47.2 | 46.31 | |
| Qwen2.5-VL-72B-InstructCoT=true, Category=Open-source Non-Reasoning Models2025.12 | 79.14 | 44.53 | 41.23 | 39.79 | |
| Qwen3-VL-235B-A22B-ThinkingParadigm=Thinking, Category=Reasoning Models2025.12 | 78.56 | 40.8 | 38.45 | 37.76 | |
| Qwen2.5-VL-72B-InstructCoT=false, Category=Open-source Non-Reasoning Models2025.12 | 78.51 | 39.93 | 37.19 | 35.76 | |
| Gemini 2.5 Pro ThinkingParadigm=Thinking, Category=Reasoning Models2025.12 | 77.19 | 53.04 | 48.58 | 47.27 | |
| Qwen2.5-VL-7B-InstructCoT=true, Category=Open-source Non-Reasoning Models2025.12 | 68.1 | 32.29 | 32.12 | 32.03 | |
| Qwen2.5-VL-7B-InstructCoT=false, Category=Open-source Non-Reasoning Models2025.12 | 55.8 | 32.81 | 32.81 | 32.81 |