High-Resolution Multimodal Reasoning on HR-Bench 8K FCP
77AccuracyRTWI
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| RTWIModel Variant=Qwen3-VL Instruct, Online setting=true2026.02 | 77 | 45.8 | |
| CISCModel Variant=Qwen3-VL Instruct, Online setting=true2026.02 | 75.8 | 30 | |
| DeepconfModel Variant=Qwen3-VL Instruct, Online setting=true2026.02 | 75.5 | 32.1 | |
| Self-Cer.Model Variant=Qwen3-VL Instruct, Online setting=true2026.02 | 75.3 | 30.8 | |
| SCModel Variant=Qwen3-VL Instruct, Online setting=true2026.02 | 74.8 | — | |
| ASCModel Variant=Qwen3-VL Instruct, Online setting=true2026.02 | 74.8 | 38.2 | |
| ESCModel Variant=Qwen3-VL Instruct, Online setting=true2026.02 | 74.8 | 25.6 | |
| BaseModel Variant=Qwen3-VL Instruct, Online setting=true2026.02 | 71.3 | — | |
| RTWIModel Variant=Qwen3-VL Thinking, Online setting=true2026.02 | 62.3 | 52.3 | |
| DeepconfModel Variant=Qwen3-VL Thinking, Online setting=true2026.02 | 61.8 | 51.1 | |
| SCModel Variant=Qwen3-VL Thinking, Online setting=true2026.02 | 61 | — | |
| ESCModel Variant=Qwen3-VL Thinking, Online setting=true2026.02 | 61 | 18.5 | |
| ASCModel Variant=Qwen3-VL Thinking, Online setting=true2026.02 | 60.8 | 28.6 | |
| Self-Cer.Model Variant=Qwen3-VL Thinking, Online setting=true2026.02 | 59 | 43.5 | |
| GPT-4oOnline setting=true2026.02 | 58.5 | — | |
| DeepEyesOnline setting=true2026.02 | 58.5 | — | |
| ThymeOnline setting=true2026.02 | 57.5 | — | |
| BaseModel Variant=Qwen3-VL Thinking, Online setting=true2026.02 | 56 | — | |
| CISCModel Variant=Qwen3-VL Thinking, Online setting=true2026.02 | 56 | 38.8 |