Long Context Understanding on HELMET
68.5AccuracySynthetic Reasoning
Evaluation Results
| Method | Links | |
|---|---|---|
| Synthetic ReasoningModel family=Qwen3 VL2026.03 | 68.5 | |
| Qwen3 VL 235B A22B InstructModel family=Qwen3 VL2026.03 | 67.6 | |
| No-thinkModel family=Qwen3 VL2026.03 | 65.9 | |
| Plain DistillationModel family=Qwen3 VL2026.03 | 65.7 | |
| LongCat-Flash Exp-ChatEvaluation Mode=Chat2025.12 | 64.7 | |
| GLM 4.6Evaluation Mode=Chat2025.12 | 64.6 | |
| Qwen Thinking TracesModel family=Mistral2026.03 | 64.1 | |
| Qwen3 VL 32B InstructModel family=Qwen3 VL2026.03 | 63 | |
| LongPOModel family=Qwen3 VL2026.03 | 62.9 | |
| Synthetic ReasoningModel family=Mistral2026.03 | 62.6 | |
| DeepSeek V3.2Evaluation Mode=Chat2025.12 | 59.5 | |
| LongCat-Flash ChatEvaluation Mode=Chat2025.12 | 59.1 | |
| No-thinkModel family=Mistral2026.03 | 55.8 | |
| Plain DistillationModel family=Mistral2026.03 | 53.1 | |
| Mistral 3.1 Small 24BModel family=Mistral2026.03 | 37 |