Long-context Visual Question Answering on MMLongBench 32K
82.4AccuracyQwen3 VL 235B A22B Instruct
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3 VL 235B A22B InstructModel family=Qwen3 VL, Scale=235B A22B2026.03 | 82.4 | |
| Qwen3 VL 32B InstructModel family=Qwen3 VL, Scale=32B2026.03 | 78.9 | |
| Synthetic ReasoningModel family=Qwen3 VL, Scale=32B2026.03 | 78.6 | |
| LongPOModel family=Qwen3 VL2026.03 | 78.4 | |
| No-thinkModel family=Qwen3 VL2026.03 | 77.7 | |
| Plain DistillationModel family=Qwen3 VL2026.03 | 77 | |
| Synthetic ReasoningModel family=Mistral, Scale=24B2026.03 | 75 | |
| Qwen Thinking TracesModel family=Mistral2026.03 | 74.2 | |
| Mistral 3.1 Small 24BModel family=Mistral, Scale=24B2026.03 | 72.9 | |
| No-thinkModel family=Mistral2026.03 | 72.2 | |
| Plain DistillationModel family=Mistral2026.03 | 71.3 |