Document Understanding on DUDE
61.8AccuracyQwen3 VL
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3 VLCheckpoint=32B2026.02 | 61.8 | |
| Qwen3 VL 32B InstructModel family=Qwen3 VL2026.03 | 61.8 | |
| Qwen3 VLCheckpoint=235B A22B2026.02 | 59.1 | |
| Qwen3 VL 235B A22B InstructModel family=Qwen3 VL2026.03 | 59.1 | |
| LongPOCheckpoint=Short Stage2026.02 | 56 | |
| LongPOModel family=Qwen3 VL2026.03 | 56 | |
| No-thinkModel family=Qwen3 VL2026.03 | 55.3 | |
| Synthetic ReasoningModel family=Qwen3 VL2026.03 | 55.1 | |
| Qwen3 VL Plain DistillationCheckpoint=Short Stage2026.02 | 54.8 | |
| Plain DistillationModel family=Qwen3 VL2026.03 | 54.8 | |
| Synthetic ReasoningModel family=Mistral2026.03 | 54.1 | |
| Mistral Plain Distillation*2026.02 | 54 | |
| Plain DistillationModel family=Mistral2026.03 | 54 | |
| Qwen Thinking TracesModel family=Mistral2026.03 | 54 | |
| Mistral 3.1 SmallCheckpoint=24B2026.02 | 52.8 | |
| Mistral 3.1 Small 24BModel family=Mistral2026.03 | 52.8 | |
| GPT-4oSize=-2026.05 | 52.7 | |
| No-thinkModel family=Mistral2026.03 | 51.4 | |
| VL-Scaler-MiMOSize=7B2026.05 | 47.6 | |
| GPT-4o-miniSize=-2026.05 | 46.5 | |
| VL-ScalerSize=7B2026.05 | 45.1 | |
| Qwen2.5-VL-InstructSize=72B2026.05 | 44.5 | |
| Pixel ReasonerSize=7B2026.05 | 44.5 | |
| MiMO-VL-InstructSize=7B2026.05 | 43.1 | |
| Qwen2.5-VL-InstructSize=7B2026.05 | 41.8 | |
| DocopilotSize=8B2026.05 | 40.7 | |
| mPLUG-Owl3Size=7B2026.05 | 39.5 | |
| VL-RethinkerSize=7B2026.05 | 39.1 | |
| Llava-OVSize=7B2026.05 | 38.1 | |
| OpenVLThinkerSize=7B2026.05 | 35.7 | |
| DeepEyesSize=7B2026.05 | 35.2 | |
| R1-VLSize=7B2026.05 | 25.4 |