Long-horizon reasoning on EXPLORE-Bench Medium
66.71Sobj ScoreGLEN
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GLEN2026.07 | 66.71 | 2.73 | |
| Qwen3-VL-8B-Thinking2026.07 | 62.61 | 2.78 | |
| Gemini-3-Pro2026.07 | 60.99 | 2.74 | |
| Qwen3-VL-8B-Instruct2026.07 | 60.78 | 2.81 | |
| GPT-5.2-Chat2026.07 | 59.88 | 2.65 | |
| LLaVA-OneVision-1.5-8B2026.07 | 51.21 | 2.44 |