Long-horizon reasoning on EXPLORE-Bench Full
65.59Sobj ScoreGLEN
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GLEN2026.07 | 65.59 | 2.69 | |
| Qwen3-VL-8B-Thinking2026.07 | 62.7 | 2.8 | |
| Gemini-3-Pro2026.07 | 60.94 | 2.75 | |
| Qwen3-VL-8B-Instruct2026.07 | 60.63 | 2.82 | |
| GPT-5.2-Chat2026.07 | 59.69 | 2.67 | |
| LLaVA-OneVision-1.5-8B2026.07 | 51.87 | 2.47 |