Visual Question Answering on UR-Bench 1.0 (test)
46.88Portrait Scrolls AccuracyQwen3-235B-A22B-Instruct
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| Qwen3-235B-A22B-Instructevaluation protocol=agent framework (Ours)2025.12 | 46.88 | 35.34 | 39.44 | 30.7 | 47.4 | 40.25 | 39.97 | |
| gpt-4.1 (2025-04-14)evaluation protocol=agent framework (Ours)2025.12 | 43.75 | 33.62 | 37.22 | 34.75 | 41.67 | 38.71 | 38.2 | |
| doubao-seed-1-6-thinkingevaluation protocol=agent framework (Ours)2025.12 | 39.06 | 31.9 | 34.44 | 29.53 | 44.27 | 37.95 | 36.75 | |
| claude-sonnet-4 (2025-05-14)evaluation protocol=agent framework (Ours)2025.12 | 37.5 | 31.03 | 33.33 | 34.04 | 44.79 | 40.19 | 37.85 | |
| DeepSeek-R1evaluation protocol=agent framework (Ours)2025.12 | 35.94 | 34.48 | 35 | 31.25 | 40.62 | 36.6 | 36.06 | |
| claude-sonnet-4 (2025-05-14)evaluation protocol=end-to-end2025.12 | 34.38 | 30.17 | 31.66 | 31.97 | 14.07 | 21.73 | 25.12 | |
| gpt4oevaluation protocol=agent framework (Ours)2025.12 | 32.81 | 27.59 | 29.44 | 39.72 | 41.67 | 40.84 | 36.95 | |
| gemini-2.5-flash-thinkingevaluation protocol=agent framework (Ours)2025.12 | 31.25 | 25 | 27.22 | 25.53 | 27.08 | 26.14 | 26.69 | |
| grok-2-vision-1212evaluation protocol=end-to-end2025.12 | 26.56 | 19.83 | 22.22 | 24.49 | 15.58 | 19.39 | 20.35 | |
| qwen-2.5-vl-32bevaluation protocol=end-to-end2025.12 | 26.56 | 20.69 | 22.78 | 31.97 | 26.13 | 28.63 | 26.64 | |
| doubao-seed-1-5-vison-proevaluation protocol=end-to-end2025.12 | 23.44 | 22.61 | 22.91 | 34.69 | 23.98 | 28.57 | 26.64 | |
| gemini-2.5-flash-thinkingevaluation protocol=end-to-end2025.12 | 21.88 | 20 | 20.11 | 15.65 | 24.62 | 20.78 | 20.55 | |
| qwen-2.5-vl-72bevaluation protocol=end-to-end2025.12 | 20.31 | 21.74 | 21.23 | 33.33 | 32.14 | 32.65 | 28.76 | |
| gpt4oevaluation protocol=end-to-end2025.12 | 12.5 | 16.52 | 15.08 | 32.88 | 18.59 | 24.71 | 21.43 | |
| o3evaluation protocol=end-to-end2025.12 | 12.5 | 24.35 | 20.11 | 15.28 | 21.32 | 18.73 | 19.2 |