Document Parsing Quality Assessment on DOCRcaseBench Text 1.0 (test)
96.43Case F1DOCR-Inspector-7B
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| DOCR-Inspector-7BCategory=Ours2025.12 | 96.43 | 81.06 | 80.21 | 81.03 | |
| Gemini 2.5 Pro ThinkingParadigm=Thinking, Category=Reasoning Models2025.12 | 88.46 | 47.17 | 32.9 | 28.16 | |
| Gemini 2.5 FlashCoT=false, Category=Proprietary Non-Reasoning Models2025.12 | 84.89 | 43.29 | 29.88 | 25.43 | |
| Gemini 2.5 FlashCoT=true, Category=Proprietary Non-Reasoning Models2025.12 | 84.75 | 42.24 | 29.74 | 25.69 | |
| Qwen3-VL-235B-A22B-ThinkingParadigm=Thinking, Category=Reasoning Models2025.12 | 83.9 | 42.02 | 31.19 | 27.46 | |
| Qwen2.5-VL-72B-InstructCoT=false, Category=Open-source Non-Reasoning Models2025.12 | 82.68 | 28.49 | 24.74 | 23.43 | |
| GPT-4oCoT=true, Category=Proprietary Non-Reasoning Models2025.12 | 77.69 | 30.54 | 27.25 | 26.35 | |
| Qwen2.5-VL-72B-InstructCoT=true, Category=Open-source Non-Reasoning Models2025.12 | 74.55 | 30.97 | 26.23 | 24.56 | |
| GPT-4oCoT=false, Category=Proprietary Non-Reasoning Models2025.12 | 72.05 | 31.66 | 28.8 | 28.04 | |
| Qwen2.5-VL-7B-InstructCoT=false, Category=Open-source Non-Reasoning Models2025.12 | 46.15 | 12.28 | 11.98 | 11.83 | |
| Qwen2.5-VL-7B-InstructCoT=true, Category=Open-source Non-Reasoning Models2025.12 | 38.17 | 12.05 | 11.64 | 11.5 |