Long-context Visual Question Answering on MMLongBench 128K
78.6AccuracyQwen3 VL 235B A22B Instruct
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3 VL 235B A22B InstructModel family=Qwen3 VL, Scale=235B A22B2026.03 | 78.6 | |
| Synthetic ReasoningModel family=Qwen3 VL, Scale=32B2026.03 | 75.7 | |
| LongPOModel family=Qwen3 VL2026.03 | 75.6 | |
| Synthetic ReasoningModel family=Mistral, Scale=24B2026.03 | 75.4 | |
| Plain DistillationModel family=Qwen3 VL2026.03 | 73.8 | |
| No-thinkModel family=Qwen3 VL2026.03 | 72 | |
| Qwen3 VL 32B InstructModel family=Qwen3 VL, Scale=32B2026.03 | 70.4 | |
| Mistral 3.1 Small 24BModel family=Mistral, Scale=24B2026.03 | 66.4 | |
| Plain DistillationModel family=Mistral2026.03 | 65.7 | |
| Qwen Thinking TracesModel family=Mistral2026.03 | 60.6 | |
| No-thinkModel family=Mistral2026.03 | 49.2 |