Long-horizon reasoning on EXPLORE-Bench Long
59.37Sobj ScoreGLEN
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GLEN2026.07 | 59.37 | 2.67 | |
| Gemini-3-Pro2026.07 | 59.17 | 2.7 | |
| GPT-5.2-Chat2026.07 | 58.06 | 2.61 | |
| Qwen3-VL-8B-Thinking2026.07 | 58.02 | 2.63 | |
| Qwen3-VL-8B-Instruct2026.07 | 56.83 | 2.71 | |
| LLaVA-OneVision-1.5-8B2026.07 | 47.62 | 2.41 |